Method and device for dividing evolutionary lineage of pathogen, electronic equipment and storage medium
By constructing a weighted monomorphic network and a modularity optimization algorithm, the problems of low automation and instability in pathogen evolutionary lineage classification are solved, achieving efficient and automated lineage classification, improving the temporal stability and biological rationality of lineage classification, and accurately reflecting the population genetic and transmission network characteristics of pathogens.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING INSTITUTE OF GENOMICS CHINESE ACADEMY OF SCIENCES (CHINA NATIONAL CENTER FOR BIOINFORMATION)
- Filing Date
- 2026-02-26
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies have low automation in pathogen evolutionary lineage classification, rely heavily on human intervention, and phylogenetic tree-based methods suffer from instability and computational scalability bottlenecks, making it difficult to accurately characterize the population genetic and transmission network characteristics of pathogens.
A weighted monomorphic network and a modularity optimization algorithm are used for community structure detection. By constructing a weighted monomorphic network and using the modularity optimization algorithm for community structure detection, combined with fine-grained modification, pathogen evolutionary lineages can be automatically identified and classified.
It achieves efficient and automated pathogen evolutionary lineage classification, improves the temporal stability and computational feasibility of the lineage classification results, accurately reflects the population genetic and transmission network characteristics of pathogens, and optimizes biological rationality.
Smart Images

Figure CN122135779A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of classification technology, and in particular to a method, apparatus, electronic device and storage medium for classifying pathogen evolutionary lineages. Background Technology
[0002] In the field of pathogen taxonomy and epidemiological surveillance, especially of viruses, one of the core tasks is to accurately classify constantly mutating pathogens into different evolutionary lineages. Traditionally, this task relies on phylogenetic tree-based analysis methods. The typical technical approach is as follows: first, a binary branching tree reflecting the evolutionary relationships of the pathogen is constructed using its genomic sequence; then, relying on expert experience or simple rules, different branches are identified and defined on this binary branching tree, and designated as different evolutionary lineages.
[0003] However, this phylogenetic lineage classification method based on phylogenetic trees has revealed significant limitations in practical applications, especially when dealing with large-scale, continuously growing pathogen genome sequences.
[0004] First, its automation level is low and it relies heavily on human intervention. For example, the widely used Pangolin system requires a large amount of manual review and subjective judgment in its genealogy definition and division process, making it difficult to adapt to the needs of automated processing of massive amounts of data, and the results are prone to lag.
[0005] Secondly, phylogenetic tree-based methods suffer from inherent instability and computational scalability bottlenecks. Because the topology of a phylogenetic tree is extremely sensitive to newly added pathogen genome sequences, the tree structure is prone to rearrangement as pathogen genome sequences are continuously updated. This can lead to non-biologically significant changes in previously defined phylogenetic classifications, disrupting the temporal continuity of long-term monitoring. Furthermore, constructing large-scale phylogenetic trees involves extremely high computational complexity.
[0006] Finally, binary branching tree models are insufficient to accurately characterize the network-like transmission characteristics of pathogens in the real world. The spread and mutation of viruses between hosts often form complex network relationships, and a strict tree structure cannot realistically and comprehensively depict these population genetic characteristics, resulting in insufficient accuracy in analyzing transmission chains. Summary of the Invention
[0007] In view of this, the purpose of this application is to provide a method, device, electronic device and storage medium for classifying pathogen evolutionary lineages, so as to realize automated lineage classification, overcome the dependence on human intervention, improve the temporal stability and computational feasibility of lineage classification results, more realistically reflect the population genetic and transmission network characteristics of pathogens, and optimize the biological rationality of automated classification.
[0008] In a first aspect, embodiments of this application provide a method for classifying pathogen evolutionary lineages, including: Obtain multiple genome sequences of the pathogen; Based on the mutation characteristics of the genome sequences, a weighted haplotype network is constructed; wherein, the nodes in the weighted haplotype network are haplotypes, and each haplotype represents a group of genome sequences with the same mutation characteristics; the weights of the edges between nodes are calculated based on the genetic similarity between the corresponding haplotypes. The community structure of the weighted monomer network is detected based on the modularity optimization algorithm, and the monomers in the weighted monomer network are divided into multiple coarse-grained clusters to obtain the coarse-grained partitioning result. The coarse-grained classification results are then refined to obtain the final pathogen evolutionary lineage classification results; wherein, the fine-grained refinement includes: In the weighted monomer network, monomers that are connected to only one downstream monomer and are classified into different coarse-grained clusters with that downstream monomer are reassigned to the cluster with the highest node-cluster similarity. And / or, For coarse-grained clusters where the number of genomic sequences corresponding to haplotypes is less than a preset threshold, the mutation similarity between the coarse-grained cluster and its parent cluster is calculated. If the mutation similarity is greater than a preset merging threshold, the coarse-grained cluster is merged into its parent cluster.
[0009] Secondly, embodiments of this application also provide a device for classifying pathogen evolutionary lineages, comprising: The acquisition module is used to acquire multiple genomic sequences of the pathogen; The first construction module is used to construct a weighted haplotype network based on the mutation characteristics of the genome sequence; wherein, the nodes in the weighted haplotype network are haplotypes, and each haplotype represents a group of genome sequences with the same mutation characteristics; the weight of the edges between nodes is calculated based on the genetic similarity between the corresponding haplotypes; The first partitioning module is used to perform community structure detection on the weighted monomer network based on the modularity optimization algorithm, and to divide the monomers in the weighted monomer network into multiple coarse-grained clusters to obtain coarse-grained partitioning results. The second partitioning module is used to refine the coarse-grained partitioning results to obtain the final pathogen evolutionary lineage partitioning results. Specifically, the second partitioning module is used to perform the fine-grained modification in the following manner: In the weighted monomer network, monomers that are connected to only one downstream monomer and are classified into different coarse-grained clusters with that downstream monomer are reassigned to the cluster with the highest node-cluster similarity. And / or, For coarse-grained clusters where the number of genomic sequences corresponding to haplotypes is less than a preset threshold, the mutation similarity between the coarse-grained cluster and its parent cluster is calculated. If the mutation similarity is greater than a preset merging threshold, the coarse-grained cluster is merged into its parent cluster.
[0010] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps in any of the possible implementations of the first aspect described above are performed.
[0011] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps in any of the possible implementations of the first aspect described above.
[0012] This application provides a method, apparatus, electronic device, and storage medium for classifying pathogen evolutionary lineages. The method uses a modularity optimization algorithm to detect community structure in weighted haplotype networks, achieving automated identification and classification of coarse-grained clusters within these networks. The entire process eliminates the need for experts to manually define branches on the tree, enabling efficient and objective processing of massive genomic sequences. This avoids the subjectivity and lag of existing technologies, achieving highly automated lineage classification and overcoming reliance on human intervention.
[0013] Furthermore, the core foundation of this embodiment is a weighted haplotype network, which, compared to a binary phylogenetic tree, is less sensitive to new data. Community detection is performed using a modularity optimization algorithm, an efficient graph clustering method, whose computational complexity is more scalable than constructing a large-scale phylogenetic tree. Therefore, this embodiment can provide a more stable phylogenetic benchmark even with dynamically growing data, meeting the technical requirements of long-term, continuous pathogen monitoring and significantly improving the temporal stability and computational feasibility of phylogenetic results.
[0014] Furthermore, the construction of a weighted haplotype network, with edge weights calculated based on the genetic similarity between haplotype nodes, can quantify the complex relationships between pathogen genome sequences. This network model itself does not require a binary branching structure, thus better accommodating and representing the complex connections that may exist in viral evolution (such as multiple ancestral contributions, recombination signals, etc.), making the final evolutionary lineage more genetically consistent and more accurately depicting the potential transmission network, thereby more realistically reflecting the population genetic and transmission network characteristics of pathogens.
[0015] Furthermore, fine-grained refinement further optimizes the biological rationality of the automated partitioning. This refinement process includes the redistribution of boundary nodes (single nodes with only one downstream connection that are fragmented) and the merging of extremely small clusters. This process effectively corrects the biologically unreasonable edge partitioning or over-fragmentation results that may be produced by purely mathematical clustering algorithms, making the final pathogen evolutionary lineage partitioning results not only mathematically compact but also more coherent and accurate in biological evolutionary logic, thus improving the precision of the lineage definition.
[0016] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart illustrating a method for classifying pathogen evolutionary lineages provided in an embodiment of this application is shown; Figure 2 This document illustrates an overall flowchart of a pathogen evolutionary lineage classification provided in an embodiment of this application. Figure 3 This illustration shows a schematic diagram of a breakpoint node reassignment provided in an embodiment of this application; Figure 4 This illustration shows a schematic diagram of an optimized merging of small-scale clusters provided in an embodiment of this application; Figure 5 This paper shows a schematic diagram of the structure of a pathogen evolutionary lineage division device provided in an embodiment of this application; Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0020] Considering that the phylogenetic lineage classification method based on phylogenetic trees has significant limitations in practical applications, especially when dealing with large-scale and continuously growing pathogen genome sequences: its automation level is low and it relies heavily on human intervention; the phylogenetic tree-based method has inherent instability and computational scalability bottlenecks; and the binary branching tree model is difficult to accurately characterize the network transmission characteristics of pathogens in the real world.
[0021] Based on this, embodiments of this application provide a method, apparatus, electronic device, and storage medium for classifying pathogen evolutionary lineages, which are described below through embodiments.
[0022] To facilitate understanding of this embodiment, a method for classifying pathogen evolutionary lineages disclosed in this application will first be described in detail. For example... Figure 1 As shown, the process includes the following steps S101-S104: S101: Obtain multiple genome sequences of the pathogen.
[0023] In this embodiment, pathogens specifically refer to viruses that can cause disease, such as SARS-CoV-2 and monkeypox virus. Each pathogen can be sequenced to obtain a complete genome sequence (i.e., a whole genome sequence), and multiple genome sequences can be obtained for each pathogen.
[0024] S102: Construct a weighted haplotype network based on the mutation characteristics of genome sequences; where nodes in the weighted haplotype network are haplotypes, and each haplotype represents a group of genome sequences with the same mutation characteristics; the weight of the edges between nodes is calculated based on the genetic similarity between the corresponding haplotypes.
[0025] In this embodiment, each genome sequence of the pathogen is preprocessed to define and extract key mutation features. Each pathogen has a reference genome sequence (referred to as the reference sequence), and different genome sequences of the pathogen contain mutations at different sites in the reference genome sequence. The mutation features of the genome sequence refer to the mutations in the pathogen's genome sequence relative to its reference sequence, forming a mutation set consisting of mutation sites and the resulting data.
[0026] In a weighted haplotype network, nodes are haplotypes, each representing a set of pathogen genome sequences with identical mutational characteristics. Edges between haplotypes represent possible direct evolutionary or transmission paths, and their weights are numerical values calculated based on the genetic similarity between corresponding haplotypes to quantify the tightness of the path (relationship).
[0027] In one possible implementation, such as Figure 2 As shown, when performing step S102, the specific steps S1021-S1024 can be executed as follows: S1021: For each pathogen's genome sequence, extract the mutations of that genome sequence relative to the reference sequence to form a mutation set consisting of mutation sites and the data after mutation. Pathogens with the same mutation set are grouped into the same haplotype, resulting in multiple haplotypes.
[0028] In this embodiment, the genome sequence of the pathogen obtained in step S101 is compared with its respective recognized, standard reference sequence (e.g., the Wuhan-Hu-1 sequence of SARS-CoV-2). Through comparison, genomic sites (i.e., mutation sites) where the pathogen's genome sequence has mutated relative to the reference sequence are identified, and the data after mutation at these mutation sites is recorded. The combination of "mutation site - mutated data" is considered a basic unit. All such basic units of a pathogen's genome sequence constitute a mutation set.
[0029] By traversing all genomic sequences of the pathogen, sequences with identical sets of mutations are grouped into the same population, which is defined as a haplotype. When constructing a weighted haplotype network, a node is used to represent that haplotype (i.e., that group of genomic sequences). In this way, millions of genomic sequences are merged and compressed into fewer, unique haplotypes, significantly reducing the complexity of subsequent computations. The core characteristic of each haplotype is the set of mutations shared by the genomic sequence group it represents.
[0030] S1022: Based on the mutation sets corresponding to each of the complete genome sequences of the pathogen, a haplotype network is constructed using a haplotype network building algorithm.
[0031] In this embodiment, to construct the topology of the weighted monolithic network, it is necessary to determine whether any two monoliths (such as monolith u and monolith v) are directly connected in terms of direct evolution or propagation. This embodiment can use the McAN algorithm (monolithic network construction algorithm) to automate this determination.
[0032] Specifically, the input to the McAN (Minimum-cost Arborescence Network) algorithm is: the mutation set of haplotype u and haplotype v, and the metadata corresponding to haplotype u and haplotype v.
[0033] The metadata includes at least several pieces of information, such as the pathogen's name, access ID, data source, related ID lineage (already defined, such as the existing Pangolin or Nextclade lineage), nucleotide integrity, sequence length, sequence quality, quality assessment, host, sample collection date, location, origin laboratory, submission date, submitting laboratory, creation time, and last update time.
[0034] The McAN algorithm analyzes the correlation between two haplotypes on the mutation set and combines it with metadata to infer whether haplotype u could be a direct evolutionary precursor of haplotype v (or vice versa), that is, to determine whether there is a connection between the two that represents a direct evolutionary or propagational relationship.
[0035] The output of the McAN algorithm is a binary adjacency indicator. :
[0036] That is, if it is determined that there is a connection between monomer type u and monomer type v, then = 1; otherwise, = 0. The relationship between all monomer pairs This forms the basis of weighted monolithic network connectivity.
[0037] S1023: For any two haplotypes in a haplotype network that have a connection relationship, calculate the Jaccard coefficient as the genetic similarity between the two haplotypes based on the mutation sets of the two haplotypes.
[0038] In this embodiment, for those determined to have a connection relationship in step S1022 (i.e. For any two haplotypes u and v with a genetic similarity of 1), it is necessary to further quantify the degree of genetic similarity between them. This embodiment uses the following Jaccard coefficient calculation formula to calculate the genetic similarity between two haplotypes;
[0039] in, and These represent the mutation sets of haplotype u and haplotype v, respectively; This indicates the genetic similarity between haplotype u and haplotype v; The value ranges from 0 to 1; the higher the genetic similarity value, the higher the proportion of mutations shared by the two haplotypes, and the more genetically similar they are.
[0040] A key advantage of this method for calculating genetic similarity is that it is insensitive to the order of magnitude of elements within the mutation set, thus effectively mitigating the impact of imbalanced data on similarity assessment.
[0041] S1024: Based on the connection relationship and genetic similarity between the two haplotypes, calculate and assign weights to the edges between the two haplotypes in the haplotype network to obtain a weighted haplotype network; wherein, the genetic similarity between haplotypes is positively correlated with the weights.
[0042] In this embodiment, the connection relationship between monomer type u and monomer type v is determined ( ) and genetic similarity ( After that, it is necessary to assign a weight with clear biological meaning to the edge between monomer u and monomer v in the weighted monomer network. .
[0043] In this embodiment, the weight of the edge between two singleton types u and v is calculated using the following exponential weighting function. :
[0044] The design of this exponential weighting formula ensures that the higher the genetic similarity between adjacent nodes (i.e., connected nodes), the greater their weight, thus giving the weighted monomorphic network structure a clear biological meaning.
[0045] Finally, based on the following adjacency matrix A weighted singularity network G is generated to provide a data foundation for subsequent cluster analysis.
[0046]
[0047]
[0048] Where V represents a singleton in a weighted singleton network, and E represents the physical adjacency relationship between two singletons. (Determined by the McAN algorithm); A represents the degree of intimacy between each pair of monotypes (weight, calculated using the exponential Jaccard coefficient).
[0049] S103: Community structure detection is performed on weighted monomer networks based on modularity optimization algorithm, and monomers in the weighted monomer network are divided into multiple coarse-grained clusters to obtain coarse-grained partitioning results.
[0050] In this embodiment, the aim is to automatically identify haplotype communities with tight internal connections and relatively sparse external connections from the weighted haplotype network constructed in step S102. Each such haplotype community is considered a candidate coarse-grained cluster. This achieves a preliminary division from a genetic relationship network to a discrete evolutionary lineage.
[0051] This embodiment employs the Leiden algorithm (an efficient modularity optimization algorithm) to detect community structure in weighted monolithic networks. The core of the Leiden algorithm is the iterative optimization of the modularity (Q) quantifier to find the optimal partitioning of the weighted monolithic network. The functional definition of modularity Q is as follows:
[0052]
[0053] in, The weight (genetic similarity) of the edge between singletons u and v in the adjacency matrix. It is the weighted degree of the singleton u (i.e., the sum of the weights of all edges connected to the singleton u). ; It is the weighted degree of the singleton v (i.e., the sum of the weights of all edges connected to the singleton v). ; m is the sum of the weights of all edges in the weighted monolithic network. ; This is an indicator function (1 when the monomers u and v belong to the same family, 0 otherwise).
[0054] A higher modularity Q value indicates a more tightly connected phylogenetic cluster. The larger the Q-value, the better the partition quality, and the sparser the inter-cluster connections. Leiden's algorithm finds the optimal partition by maximizing the Q-value.
[0055] The Leiden algorithm iteratively optimizes the modularity Q by locally moving nodes to increase the connection density within communities while minimizing the connection distance between communities. During optimization, the Leiden algorithm not only focuses on the local affiliation of nodes but also automatically eliminates and merges disconnected communities to ensure topological connectivity for each partitioned community. As the algorithm progresses, the network is compressed into a macroscopic structure where each node represents a unique community. Local node movements then continue at this new structural level to further optimize the modularity Q. This iterative process continues until the modularity function Q reaches its maximum value and no longer shows significant improvement; the community structure at this point represents the optimal coarse-grained partitioning of the network.
[0056] In one possible implementation, such as Figure 2 As shown, when executing step S103, the specific steps S1031-S1033 can be performed as follows: S1031: Initialize each monomer in the weighted monomer network as the initial cluster center of an independent cluster.
[0057] In this embodiment, each monomer in the weighted monomer network is initialized as the initial cluster center of an independent cluster. That is, in the initial state, each monomer forms its own cluster, and the modularity Q is usually not optimal.
[0058] S1032: Execute the iterative process, which includes a local movement phase and a network compression phase; During the local movement phase, for each monomer in the current monomer network, the modularity gain caused by moving to any adjacent cluster is calculated. If there is a positive modularity gain, it is moved to the adjacent cluster that brings the maximum modularity gain; if there is no movement that brings a positive modularity gain, its current cluster affiliation remains unchanged. During the network compression phase, all monomers belonging to the same cluster after local movement are compressed into a new supernode, and a new compressed network is constructed based on the original weight relationship. Treat the compressed network as a new weighted single-unit network and repeat the local movement phase and the network compression phase. The iteration stops when a positive modularity gain cannot be obtained through node movement in a weighted monolithic network.
[0059] In this embodiment, an iterative optimization process is performed. The purpose of this iterative process is to gradually increase the modularity Q. It iteratively executes two core phases: a local migration phase and a network compression phase, until the convergence condition is met.
[0060] During the local move phase, the Leiden algorithm traverses every monomer in the current weighted monomer network. For each monomer, the Leiden algorithm calculates the change in modularity caused by moving it to any of its neighboring clusters (i.e., clusters directly connected to the monomer by edges). This change is called the modularity gain. The Leiden algorithm identifies the neighboring cluster that provides the largest positive modularity gain among all possible moves. If such a positive modularity gain exists, the monomer is moved to that neighboring cluster; otherwise, if no move provides a positive modularity gain (i.e., all possible moves would decrease or fail to increase modularity), the monomer remains in its current cluster. This process, through repeated local adjustments, enhances the connection density within each cluster.
[0061] During the network compression phase, after one round of local moves, the weighted monomer network is divided into several clusters. In this phase, the Leiden algorithm compresses all monomers belonging to the same cluster into a new supernode.
[0062] Then, based on the connectivity between monomers in the original weighted monomer network, the connection weights between these supernodes are recalculated (for example, the weights of all edges between monomers originally belonging to two different clusters are summed to serve as the edge weights between the new supernodes), thereby constructing a new, smaller compressed network. This compressed network reflects the macroscopic relationships between clusters after the previous partitioning.
[0063] After completing one local move and network compression, the resulting compressed network is used as a new weighted monolithic network, and the above two stages are repeated. This iteration allows the Leiden algorithm to continuously optimize the modularity.
[0064] When the Leiden algorithm, after scanning the local move phase in a new weighted monomorphic network (which may be the original weighted monomorphic network or a compressed network after some compression), fails to find a move that can generate a positive modularity gain for any monomorph, it indicates that the modularity Q has reached a local optimum, the Leiden algorithm converges, and the iteration stops.
[0065] S1033: The cluster partitioning result obtained after the final iteration is used as the coarse-grained partitioning result.
[0066] In this embodiment, once the iteration process meets the stopping condition, the Leiden algorithm outputs the final coarse-grained partitioning result. This result reflects the optimized community structure at multiple levels, with the haplotypes within each coarse-grained cluster (which may correspond to the original haplotype or a supernode at a certain level) being highly correlated in genetic relationships. This coarse-grained partitioning result represents multiple coarse-grained clusters automatically partitioned from the weighted haplotype network, providing a foundation for the next step of fine-tuning.
[0067] S104: Fine-grained modification of the coarse-grained classification results to obtain the final pathogen evolutionary lineage classification results; Fine-grained modification includes: In the weighted monomer network, monomers that are connected to only one downstream monomer and are classified into different coarse-grained clusters with that downstream monomer are reassigned to the cluster with the highest node-cluster similarity. And / or, For coarse-grained clusters where the number of genomic sequences corresponding to haplotypes is less than a preset threshold, the mutation similarity between the coarse-grained cluster and its parent cluster is calculated. If the mutation similarity is greater than a preset merging threshold, the coarse-grained cluster is merged into its parent cluster.
[0068] In this embodiment, the coarse-grained division results obtained in step S103 are finely adjusted in a biological sense to correct edge misjudgment and fragmentation that may occur in pure topological clustering, thereby improving the accuracy and biological consistency of phylogenetic division.
[0069] Among them, such as Figure 2 As shown, fine-grained modification mainly includes two operations: the first is the redistribution of "breakpoint nodes" (corresponding to steps S10401-S10405); the second is the merging of "small-scale clusters" (corresponding to steps S10406-S10410). This section first details the specific implementation method of breakpoint node redistribution.
[0070] The purpose of breakpoint node reassignment is to address the problem of key nodes (called breakpoint nodes) located at lineage boundaries and connected to only one downstream haplotype node being incorrectly classified during community structure detection. By calculating the genetic similarity between the breakpoint node and its current coarse-grained cluster as well as the coarse-grained cluster to which its downstream haplotype node belongs, and reassigning it based on the similarity, its affiliation can be made more consistent with evolutionary continuity.
[0071] In one possible implementation, when performing step S104, which involves reassigning a single-monotype in the weighted monotype network that is connected to only one downstream monotype and is classified into a different coarse-grained cluster with that downstream monotype to the cluster with the highest node-cluster similarity, the specific steps S10401-S10405 can be followed: S10401: Identify monomer types from each coarse-grained cluster as nodes to be reallocated, provided that they meet the following conditions: In the weighted monomer network, the monomer type has one and only one directly connected downstream monomer type, and the monomer type and its downstream monomer type are located in different coarse-grained clusters.
[0072] In this embodiment, from the multiple coarse-grained clusters obtained in step S103, all monomer types are traversed, and monomer types that meet the following two conditions are identified as nodes to be reallocated (i.e., breakpoint nodes): 1) In a weighted monomorphic network, the monomorphic entity has one and only one directly connected downstream monomorphic entity. Here, "downstream" usually refers to a successor node that appears later in the evolutionary or propagation timeline, or is defined according to the directionality of the weighted monomorphic network.
[0073] 2) The node to be reassigned and its only downstream monomer are located in different coarse-grained clusters in the current coarse-grained partitioning results.
[0074] Nodes that meet both of these conditions and are to be reassigned are prone to bias in their assignment due to the hard boundaries of mathematical clustering algorithms, and therefore require special handling.
[0075] For example, such as Figure 3 As shown, each red dot represents a single entity; the purple solid circle represents a coarse-grained cluster C1 before fine-grained modification, and the green solid circle represents a coarse-grained cluster C2 before fine-grained modification; the red dots circled in the yellow dashed box are nodes to be reallocated that simultaneously meet two conditions.
[0076] The purple dashed circle represents a cluster (i.e., an evolutionary lineage) C1' after fine-grained modification, and the green dashed circle represents a cluster (i.e., an evolutionary lineage) C2' after fine-grained modification.
[0077] S10402: For each node to be reassigned, calculate the first Jaccard coefficient between the node to be reassigned and its current first coarse-grained cluster; wherein, the first Jaccard coefficient is calculated based on the mutation set of the node to be reassigned and the mutation set shared by all haplotypes in the first coarse-grained cluster; the first Jaccard coefficient is used to represent the node-cluster similarity between the node to be reassigned and its current first coarse-grained cluster.
[0078] In this embodiment, for each identified node to be reassigned, the genetic similarity between the node to be reassigned and its current coarse-grained cluster (denoted as the first coarse-grained cluster) is first calculated.
[0079] This embodiment uses the Node-Cluster Jaccard Coefficient (NC-JC) for measurement. Specifically, the first Jaccard coefficient between the node to be reassigned and its current first coarse-grained cluster is calculated using the following formula:
[0080] Where u is the node to be reallocated. C1 represents the set of mutations of node u to be reallocated; C1 represents the first coarse-grained cluster to which node u currently belongs. It is the set of mutations common to all haplotypes within the first coarse-grained cluster; Let be the first Jaccard coefficient between the node u to be reallocated and the first coarse-grained cluster C1.
[0081] S10403: Calculate the second Jaccard coefficient between the node to be reassigned and the second coarse-grained cluster to which its downstream haplotype belongs; wherein, the second Jaccard coefficient is calculated based on the mutation set of the node to be reassigned and the mutation set shared by all haplotypes in the second coarse-grained cluster; the second Jaccard coefficient is used to represent the node-cluster similarity between the node to be reassigned and the second coarse-grained cluster to which its downstream haplotype belongs.
[0082] In this embodiment, the genetic similarity between the node to be reassigned u and the coarse-grained cluster (denoted as the second coarse-grained cluster C2) to which its downstream haplotype belongs is calculated. The node-cluster Jaccard coefficient is also used. Specifically, the second Jaccard coefficient between the node to be reassigned and the second coarse-grained cluster to which its downstream haplotype belongs is calculated using the following formula:
[0083] Where u is the node to be reallocated. C1 represents the set of mutations for node u to be reallocated; C2 represents the second coarse-grained cluster to which the downstream haplotype of node u belongs. It is the set of mutations common to all haplotypes within the second coarse-grained cluster; Let be the second Jaccard coefficient between the node u to be reallocated and the second coarse-grained cluster C2.
[0084] S10404: Compare the first Jaccard coefficient with the second Jaccard coefficient.
[0085] In this embodiment, the calculated first Jaccard coefficient With the second Jaccard coefficient Compare them.
[0086] S10405: If the second Jaccard coefficient is greater than the first Jaccard coefficient, then the node to be reassigned is reassigned from the first coarse-grained cluster to the second coarse-grained cluster; if the second Jaccard coefficient is less than or equal to the first Jaccard coefficient, then the node to be reassigned remains unchanged in the first coarse-grained cluster.
[0087] In this embodiment, the node affiliation is reallocated based on the comparison results: like > If the genetic similarity between the node to be reassigned and the second coarse-grained cluster C2 to which the downstream haplotype belongs is higher than its similarity with the current first coarse-grained cluster C1, then the node to be reassigned is determined to be genetically closer to the second coarse-grained cluster C2. Therefore, the node to be reassigned is removed from the current first coarse-grained cluster C1 and reassigned to the second coarse-grained cluster C2.
[0088] like ≤ If the similarity between the node u to be reassigned and the second coarse-grained cluster C2 is not higher than that with the first coarse-grained cluster C1, then the current partitioning is considered reasonable or requires no adjustment. Therefore, the node u to be reassigned remains unchanged in its affiliation with the first coarse-grained cluster C1.
[0089] The above steps correct the potential misclassification of boundary nodes in the coarse-grained classification results, making the phylogenetic classification at the boundaries more accurate and more in line with the evolutionary continuity in biology.
[0090] After completing the redistribution of boundary breakpoint nodes (nodes to be reallocated) (S10401-S10405), the next step in fine-grained modification is to optimize and merge small clusters. During the community structure detection process in step S103, some coarse-grained clusters with very few haplotypes (e.g., fewer than 10 haplotypes) may be generated. These coarse-grained clusters may be caused by sequencing noise, rare mutations, or small branches generated by the algorithm in complex network topologies. Retaining too many such small clusters will increase the complexity of subsequent analysis and may lead to fragmented results, affecting the clarity and statistical significance of lineage definitions. Therefore, this step aims to merge eligible small clusters back into their parent clusters with clear evolutionary origins based on genetic similarity.
[0091] In one possible implementation, in step S104, for coarse-grained clusters containing haplotypes whose number of corresponding genomic sequences is less than a preset threshold, the mutation similarity between the coarse-grained cluster and its parent cluster is calculated. If the mutation similarity is greater than a preset merging threshold, the coarse-grained cluster is merged into its parent cluster. Specifically, this can be performed according to the following steps S10406-S10412: S10406: Identify small-scale clusters from each coarse-grained cluster where the number of genomic sequences corresponding to the haplotypes contained is less than a preset threshold.
[0092] In this embodiment, from the multiple coarse-grained clusters processed as described above, all coarse-grained clusters containing fewer than a preset number threshold are identified. The preset number threshold is a configurable parameter, for example, it can be set to 10. Coarse-grained clusters that meet this condition are defined as small-scale clusters and are candidates for merging and evaluation.
[0093] S10407: For each small-scale cluster, determine the coarse-grained cluster that has a connection with it in the weighted monomorphic network as its parent cluster.
[0094] In this embodiment, for each small-scale cluster identified in S10406, its parent cluster needs to be determined within the weighted monolithic network. The parent cluster refers to a larger, coarse-grained cluster that has a direct and strong connection to the small-scale cluster in the weighted monolithic network topology. Typically, this can be determined by finding the neighboring coarse-grained cluster with the largest sum of edge weights or the most connections to the small-scale cluster. This parent cluster is most likely the direct source or primary associated group of the small-scale cluster in terms of evolutionary relationships.
[0095] For example, such as Figure 4 As shown, each red dot represents a monomer, and the red dots circled in the yellow dashed box represent the identified small clusters. In this example, there are two small clusters. The green dashed circle represents the parent cluster of the small clusters.
[0096] S10408: Determine the common mutation set of small-scale clusters; the common mutation set consists of mutation sites and data after mutation that are shared by monomers of no less than a preset proportion within the cluster.
[0097] In this embodiment, the common mutation set of the small-scale cluster is calculated. The common mutation set of the small-scale cluster consists of mutation sites and mutated data shared by at least a preset proportion of monomers within the small-scale cluster.
[0098] This set aims to identify mutations shared by most members within a small cluster that represent its core genetic characteristics. The frequency of each "mutation site-mutated data" pair within the small cluster is counted. Mutation sites with a frequency reaching or exceeding a predetermined proportion (e.g., 60%) are included in the set. The set of these frequently shared mutation sites constitutes the common mutation set of the small cluster.
[0099] S10409: Determine the common mutation set of the parent cluster.
[0100] In this embodiment, the common mutation set of the parent cluster consists of the mutation sites and the data after mutation that are common to the monomers in the parent cluster at a rate of not less than a preset proportion.
[0101] Using the same method as in step S10408, calculate the common mutation set of the parent cluster. That is, identify the mutation set shared by haplotypes exceeding a predetermined proportion (e.g., 60%) within the parent cluster. This set represents the core genetic marker of the parent cluster.
[0102] S10410: Based on the shared mutation set of small-scale clusters and the shared mutation set of the parent cluster, calculate the cluster-cluster Jaccard coefficient as the mutation similarity between the small-scale cluster and the parent cluster.
[0103] In this embodiment, to quantitatively assess the similarity of small clusters with their parent clusters in core genetic features, the cluster-cluster Jaccard coefficient (CC-JC) is used as an indicator of mutational similarity. Specifically, the mutational similarity (cluster-cluster Jaccard coefficient) between the small cluster and the parent cluster is calculated using the following formula:
[0104] in, For small-scale clusters Cs, this is the common mutation set; The set of common mutations of the parent cluster Cp. The mutation similarity between small clusters and their parent clusters (cluster-cluster Jaccard coefficient).
[0105] S10411: Compare mutation similarity with a preset merging threshold.
[0106] In this embodiment, the calculated mutation similarity (cluster-cluster Jaccard coefficient) will be used. The similarity is compared with a preset merging threshold. The merging threshold is a configurable parameter used to set a similarity threshold for determining whether a small cluster should be merged with its parent cluster. In a preferred embodiment, the merging threshold can be set to 0.6.
[0107] S10412: If the mutation similarity is greater than the merging threshold, the small clusters are merged into the parent cluster; if the mutation similarity is less than or equal to the merging threshold, the small clusters remain independent.
[0108] In this embodiment, if mutation similarity If the merging threshold is >0.6, the small cluster is determined to be highly similar to its parent cluster in core genetic characteristics and lacks significant uniqueness. Therefore, it is merged into the parent cluster. After merging, all haplotypes in the original small cluster become members of its parent cluster, and the original small cluster no longer exists as an independent cluster.
[0109] For example, such as Figure 4 As shown, the solid green circle represents merging two small clusters into the parent cluster.
[0110] If mutation similarity If the value is ≤ merging threshold (e.g., ≤ 0.6), it is determined that even though the small cluster is small, its core mutational features are sufficiently different from those of its parent cluster, potentially representing a new and promising microevolutionary direction. Therefore, the small cluster is kept independent and not merged.
[0111] By sequentially performing two fine-grained modification operations—breakpoint node reallocation (S10401-S10405) and small-scale cluster merging (S10406-S10412)—step S104 ultimately outputs a pathogen evolutionary lineage classification result that is topologically robust, genetically consistent, and biologically meaningful. Compared to the coarse-grained classification result, this result has higher classification accuracy and interpretability.
[0112] This embodiment not only applies to the phylogenetic classification of static genome sequences but also incorporates a dynamic update mechanism to address the reality that pathogen genome sequences continue to grow over time. This embodiment ensures that existing phylogenetic classification results can be incrementally updated after obtaining new pathogen genome sequences, while maximizing the historical continuity and temporal stability of phylogenetic definitions, thus avoiding the phylogenetic classification fluctuations (instability) caused by data updates in traditional methods.
[0113] In one possible implementation, such as Figure 2 As shown, after obtaining the final pathogen evolutionary lineage classification result output in step S104 (referred to as the current lineage set), if a new genome sequence is obtained, dynamic updates can be performed according to the following steps S1051-S1053: S1051: The genome sequence of the newly added weighted haplotype network is used as incremental data. The structure of the weighted haplotype network is updated based on the newly added genome sequence. Community structure detection and fine-grained modification are re-performed on the updated weighted haplotype network to generate a candidate new lineage set independent of the current lineage set. The current lineage set is the pathogen evolution lineage division result before the addition of incremental data. The candidate new lineage set consists of multiple candidate new evolutionary lineages.
[0114] In this embodiment, the newly acquired (newly added) genome sequences are used as incremental data and merged with historical data (i.e., all genome sequences constituting the current weighted haplotype network). Based on the merged full genome sequence data, step S102 is repeated, that is, updating the structure of the weighted haplotype network (including adding haplotypes, updating the connection relationships and weights between haplotypes).
[0115] Then, the community structure detection in step S103 and the fine-grained refinement in step S104 are re-executed on the updated weighted haplotype network. This recalculation process will produce a new, independent lineage classification result, which is defined as the candidate new lineage set. The current lineage set is the pathogen evolutionary lineage classification result before the addition of incremental data, and the current lineage set consists of multiple current evolutionary lineages. Similarly, the candidate new lineage set consists of multiple candidate new evolutionary lineages.
[0116] It should be noted that, due to the characteristics of the weighted single-type network structure and clustering algorithm, the candidate new lineage set may differ from the current lineage set in some regions.
[0117] S1052: For any current evolutionary lineage and a candidate new evolutionary lineage, calculate the aggregation ratio and dispersion ratio from the current evolutionary lineage to the candidate new evolutionary lineage; wherein, the aggregation ratio is the proportion of the number of genomic sequences corresponding to haplotypes shared by the current evolutionary lineage and the candidate new evolutionary lineage to the total number of genomic sequences corresponding to haplotypes contained in the candidate new evolutionary lineage; the dispersion ratio is the proportion of the number of genomic sequences corresponding to haplotypes shared by the current evolutionary lineage and the candidate new evolutionary lineage to the total number of genomic sequences corresponding to haplotypes contained in the current evolutionary lineage.
[0118] In this embodiment, in order to objectively evaluate the continuity and inheritance relationship of phylogenetic evolution between the current phylogenetic set and the candidate new phylogenetic set, two core quantitative indicators are introduced: Aggregated Proportion and Dispersed Proportion.
[0119] For any current evolutionary lineage (denoted as ) ) and a candidate new evolutionary lineage (denoted as ) ): Aggregation ratio: It is defined as the current evolutionary lineage. With candidate new evolutionary lineages The number of shared haplotypes, accounting for a percentage of candidate new evolutionary lineages The proportion of the total number of monomers contained.
[0120] Specifically, the aggregation ratio from the current evolutionary lineage to the candidate new evolutionary lineage is calculated using the following formula. :
[0121] in, This represents the total number of haplotypes contained in the candidate new evolutionary lineage; Indicates the current evolutionary lineage With candidate new evolutionary lineages The number of common monomer types. The higher the value (e.g., close to 1), the more likely it is to be a candidate new evolutionary lineage. Mainly composed of the current evolutionary lineage It evolved from something else, and the inheritance relationship is clear.
[0122] Dispersion ratio: It is defined as the current evolutionary lineage. With candidate new evolutionary lineages The number of shared haplotypes accounts for a certain percentage of the current evolutionary lineage. The proportion of the total number of monomers contained.
[0123] Specifically, the dispersion ratio from the current evolutionary lineage to the candidate new evolutionary lineage is calculated using the following formula. :
[0124] in, This represents the total number of haplotypes contained in the current evolutionary lineage. The higher the value, the better the current evolutionary lineage. A significant portion of the members migrated to candidate new evolutionary lineages. .
[0125] S1053: Based on the aggregation ratio and dispersion ratio, hierarchical lineage classification is performed on the genomic sequences corresponding to haplotypes in the candidate new lineage set to correct the classification of haplotypes in the candidate new lineage set, thereby generating updated pathogen evolutionary lineage classification results. The hierarchical classification of haplotypes in the candidate new lineage set includes: If a monomer type is assigned to the root lineage in the current lineage set and is classified into a non-root lineage in the candidate new lineage set, then the assignment change is accepted when the aggregation ratio from the root lineage to the non-root lineage is greater than or equal to the first preset threshold; otherwise, the monomer type is forced to remain in the root lineage after the update. If a haplotype is assigned to any first-generation lineage in the current lineage set and is assigned to the root lineage in the candidate new lineage set, then the assignment change is accepted if the dispersion ratio from the first-generation lineage to the root lineage is less than the second preset threshold; otherwise, the haplotype is forced to remain in the original first-generation lineage after the update. For haplotypes that do not belong to the root lineage in the current lineage set and whose dispersion ratio between their original lineage and the root lineage fails to meet the second preset threshold, their affiliation in the candidate new lineage set is modified to the candidate new evolutionary lineage that has the largest aggregation ratio with their original lineage in the current lineage set.
[0126] In this embodiment, directly adopting the candidate new lineage set may lead to unexpected and unstable changes in the lineage classification of some haplotypes. Therefore, this step, based on the aggregation and dispersion ratios calculated in S1052, performs a hierarchical classification logic prioritizing stability on the haplotypes in the candidate new lineage set, thereby correcting the preliminary results and generating updated pathogen evolutionary lineage classification results. The hierarchical classification rules are as follows: 1) Root lineage stability determination: If a haplotype belongs to the root lineage (representing the most basic and core evolutionary group) in the current lineage set, but is assigned to a non-root lineage in the candidate new lineage set, this change requires calculating the aggregation ratio r¹ of the haplotype's original root lineage to the non-root lineage. Only when r¹ is greater than or equal to a first preset threshold (preferably, the first preset threshold can be set to 0.6) is the new cluster (the non-root lineage) considered to have sufficiently high inheritance purity from the original root lineage members, thus accepting this assignment change. Otherwise, to maintain classification stability, the haplotype is forced to remain a root lineage member in the updated results.
[0127] 2) Determination of the survival of the first-generation lineage: If a haplotype belongs to a first-generation lineage in the current lineage set (i.e., a major branch directly derived from the root lineage), but is reclassified back to the root lineage in the candidate new lineage set, then calculate the dispersion ratio r from the original first-generation lineage to the root lineage. 0 Only when r 0 If the value is less than the second preset threshold (preferably, the second preset threshold can be set to 0.1), it indicates that the structure of the original first-generation lineage has basically dissipated, and it is accepted to be classified into the root lineage. Otherwise, it indicates that the first-generation lineage still has a significant structure, and the haplotype is forced to still be classified into the original first-generation lineage after the update.
[0128] 3) Optimization of Affiliation for Other Descendant Lineages: For haplotypes that do not fall into the above two categories (i.e., ordinary descendant lineage nodes), their affiliation correction follows the "maximum inheritance" principle. That is, the aggregation ratio r¹ between the haplotype's current evolutionary lineage in the original current lineage set and each candidate new evolutionary lineage in the candidate new lineage set is calculated. Then, in the updated result, the haplotype is corrected to the candidate new evolutionary lineage with the maximum aggregation ratio r¹ to its original current evolutionary lineage. This ensures that haplotypes are always assigned to a new group most similar to their historical origin, achieving a smooth transition.
[0129] Through the above three-layer judgment rules, the dynamic update mechanism effectively suppresses the discontinuous changes in lineage labels caused by data increment or algorithm randomness while adding new genome sequences and reflecting real evolution. Finally, it outputs a stable updated pathogen evolution lineage classification result that includes the latest genome sequences and maintains high consistency with historical classifications.
[0130] After obtaining the updated pathogen evolutionary lineage classification results through a dynamic update mechanism, this embodiment also includes a set of automated evolutionary lineage naming methods. This naming method aims to assign a unique name identifier to each lineage, which not only ensures uniqueness but also clearly reflects the evolutionary hierarchy and historical inheritance relationships between lineages, thereby supporting the management, traceability, and exchange of large-scale, long-term pathogen gene monitoring data.
[0131] In one possible implementation, such as Figure 2 As shown, after executing step S1053, the updated pathogen evolutionary lineage classification result is recorded as the newly formed lineage set, and the candidate new evolutionary lineages in the updated pathogen evolutionary lineage classification result are recorded as newly formed lineages; the method can also be executed according to the following steps S1061-S1065: S1061: Construct a cross-relationship matrix between the new lineage and the current evolutionary lineage; wherein, the rows of the cross-relationship matrix correspond to the current evolutionary lineage, the columns correspond to the new lineage, and the element values in the cross-relationship matrix are the number of genomic sequences corresponding to haplotypes shared between the lineages in the corresponding rows and columns.
[0132] In this embodiment, in order to accurately quantify the lineage evolution correspondence between the newly formed lineage set (i.e., the updated pathogen evolution lineage classification result) and the current lineage set (the pathogen evolution lineage classification result before the update, i.e., the historical classification result), a cross-relationship matrix M is constructed.
[0133] The rows of this cross-relationship matrix correspond to the m current evolutionary lineages in the current lineage set, and the columns correspond to the n newly formed lineages in the newly formed lineage set. The element value M(i,j) in the i-th row and j-th column of the cross-relationship matrix is defined as the number of genomic sequences corresponding to haplotypes shared between the i-th current evolutionary lineage and the j-th newly formed lineage. This matrix objectively and quantitatively depicts the overlap and divergence in the membership composition of old and new lineages, and is the core basis for subsequent automated naming decisions.
[0134] S1062: Traverse the cross-relation matrix. For each new lineage in the new lineage set, if the new lineage is a first-generation lineage directly derived from the root lineage, determine the current evolutionary lineage with the largest proportion of genome sequence source corresponding to its haplotype according to the cross-relation matrix, and inherit the name identifier from the current evolutionary lineage; if the name identifier is not occupied, it is assigned; otherwise, a new first-generation lineage name identifier is assigned to it.
[0135] In this embodiment, each column of the cross-relationship matrix M (i.e. each new lineage) is traversed, and it is first determined whether it is a first-generation lineage directly derived from the root lineage (which is represented in the evolutionary network as being directly connected to the root lineage representing the starting point of evolution).
[0136] Inheritance Assignment: Examine the columns corresponding to the new lineage in the cross-relationship matrix and find the row with the largest element value. The current evolutionary lineage corresponding to this row is the primary source of haplotypes for this new lineage. If this primary source current evolutionary lineage is itself a first-generation lineage, and its complete name identifier (e.g., "0112.0001") has not yet been assigned to any other new lineage, then this name identifier is directly inherited by this new lineage.
[0137] New Allocation: If inheritance is not possible (e.g., the primary source cluster is not first-generation, or its name is already taken), a new, unique first-generation lineage name identifier is automatically generated and allocated. The generation rule for the new identifier is: among all currently allocated first-generation lineage name identifiers, find the one with the largest value, and then increment it. For example, if the existing largest identifier is "0112", the newly generated one will be "0113". An initial descendant sequence number, such as "0113.0001", can be assigned to complete its full name identifier.
[0138] S1063: If the new lineage does not belong to the first generation lineage, then in the cross-relationship matrix, search for whether there exists a current evolutionary lineage such that the number of genome sequences corresponding to haplotypes shared by the new lineage and the current evolutionary lineage is simultaneously the maximum value of the row where the current evolutionary lineage is located and the maximum value of the column where the new lineage is located; if such a lineage exists, then the new lineage is determined to be the direct successor of the current evolutionary lineage, and a name identifier of the current evolutionary lineage is assigned to it.
[0139] In this embodiment, for newly formed lineages that are determined not to belong to the first generation, a stricter bidirectional maximum rule is used to determine whether they should directly inherit the name of a current evolutionary lineage. In the cross-relationship matrix M, for the j-th column where the newly formed lineage is located: Find the element with the largest value in the column, and denote its row as imax. This indicates that the current evolutionary lineage imax contributes the most haplotypes to the new lineage.
[0140] Meanwhile, the element with the largest value in the imax row is found, and its column is denoted as jmax, indicating that the new lineage jmax received the most haplotypes from the current evolutionary lineage imax.
[0141] If jmax is exactly equal to the j-th column currently being processed (i.e., jmax = j), it means that the new lineage j and the current evolutionary lineage imax have a primary flow relationship and a clear and explicit direct evolutionary inheritance relationship. In this case, the new lineage is determined to be the direct inheritor of the current evolutionary lineage imax, and the fully qualified name identifier of imax (such as "0112.0153") is directly assigned to it.
[0142] S1064: If a new lineage is neither a first-generation lineage directly derived from the root lineage nor can it be determined as a direct successor of any current evolutionary lineage, then a new incremental name identifier is assigned to it; the new incremental name identifier is generated by keeping the parent identifier of its first-generation lineage unchanged and setting the offspring sequence number within that first-generation lineage to the increment of the current maximum sequence number.
[0143] In this embodiment, if a new lineage is neither a first-generation lineage (S1062) nor can a direct successor be matched through the bidirectional maximum value rule (S1063), it is determined to be a newly generated evolutionary branch and a new incremental name identifier needs to be assigned to it.
[0144] At this point, it is necessary to first determine the family to which this new lineage (i.e., the newly generated evolutionary branch) belongs, i.e., its paternal identifier. This is usually achieved by analyzing the topology of a weighted haplotype network to determine the upstream new lineage (called its paternal lineage) that is most closely connected to the new lineage and is most likely its direct evolutionary source. The part before the "." in the name identifier of that paternal lineage is the paternal identifier that the new branch should inherit. For example, if the name of the paternal lineage is "0112.0153", then the paternal identifier of the new lineage is "0112".
[0145] After determining the parent identifier (e.g., "0112"), keep this identifier unchanged. Then, scan within this family (i.e., all new lineages starting with "0112") to find the maximum value of the offspring sequence number (the numeric part after ".") among all currently assigned name identifiers. Increment this maximum value to generate a new offspring sequence number.
[0146] The unchanged parent identifier is combined with the newly generated child sequence number to form a complete new incremental name identifier. For example, in the "0112" family, if the existing maximum child sequence number is "0250", then the newly generated branch will be named "0112.0251".
[0147] S1065: After allocating name identifiers for all new lineages, associate the finalized name identifiers with the new lineage set to output a new lineage set that has been associated with the name identifiers.
[0148] In this embodiment, after traversing all newly formed lineages and assigning name identifiers according to the above method, a final verification is performed to ensure the uniqueness of all identifiers within the entire namespace. Subsequently, each newly formed lineage is bound to its assigned name identifier. Finally, a set of newly formed lineages associated with name identifiers is output, representing a complete pathogen evolutionary lineage classification result where each lineage has a unique, standardized name reflecting its evolutionary position and historical origin. This result can be directly used to build databases and generate monitoring reports, achieving full automation and intelligence from data partitioning to standardized naming.
[0149] Next, in order to objectively and comprehensively verify the technical advantages of this embodiment in pathogen lineage classification, three core quantitative indicators were introduced for comprehensive evaluation: Weighted Information Entropy: Used to assess genetic consistency within a lineage. This metric introduces haplotype loss rate as a penalty term into the traditional Shannon entropy, and its calculation formula is defined as follows:
[0150] in, This indicates the proportion of genome sequences belonging to haplotype b in the total pedigree sequence N. This represents the deletion rate of the haplotype. The lower the weighted information entropy value, the more homogeneous the haplotype composition and the higher the genetic purity within the lineage.
[0151] The Simulated Davies-Bouldin Index (SDBI) is used to quantify the intra-cluster tightness and inter-cluster separation of clusters. This index calculates the relative distance between cluster pairs based on the Jaccard distance, and its formula is defined as follows:
[0152] in, Representative lineage With genealogy Inter-lineage distance between classes. and These represent their respective intra-lineage distances. The lower the value, the clearer the phylogenetic boundaries and the better the differentiation.
[0153] Lineage Retention Rate (LRR): This metric measures the temporal stability of a system. It calculates the proportion of sequences whose lineage remains unchanged between two consecutive data update cycles. The formula is defined as follows:
[0154] in, Indicates from time arrive The number of sequences that remain within the same lineage during the period. This represents the total number of sequences shared between the two time points. A higher lineage retention rate reflects the stronger robustness of this embodiment in dealing with dynamic data growth.
[0155] The performance of the method in this embodiment was evaluated using the above-mentioned metrics: 1) It overcomes the lag of traditional methods and achieves extremely high classification stability. Existing technologies, relying on the reconstruction of binary phylogenetic trees, often lead to drastic fluctuations in old lineage classifications with the addition of new data. This embodiment ensures classification continuity over time by introducing a dynamic update mechanism for dispersion and aggregation ratios. Experimental data shows that in continuous monitoring of SARS-CoV-2 for 48 months (January 2021 to December 2024), the average lineage retention rate (LRR) of this embodiment remained above 0.94, with nearly 90% of the months exceeding 0.90. Similarly, in validation against monkeypox virus, the lineage retention rate of this invention remained stable between 0.95 and 1.00 from January 2022 to November 2024, with an average exceeding 0.99. These results demonstrate that this embodiment effectively solves the classification instability problem caused by massive data growth, providing a reliable benchmark for long-term epidemiological monitoring.
[0156] 2) Improved genetic consistency and taxonomic purity within the lineage. This embodiment constructs a weighted haplotype network and combines it with modularity optimization to make each lineage more uniform in genetic characteristics. To quantify this effect, this embodiment introduces Weighted Information Entropy (WIE) for evaluation; the lower the entropy value, the smaller the differences within the lineage. Comparative analysis shows that on the SARS-CoV-2 dataset, the average WIE value of the lineages divided in this embodiment is 3.46, which is better than the Pangolin system's 3.68; on the monkeypox virus dataset, the average WIE value of the lineages in this invention is 3.55, which is significantly lower than the Nextclade system's 4.98. The significant left-skewed trend of the WIE distribution confirms that the lineages divided in this embodiment have higher genetic consistency and reduce classification ambiguity.
[0157] 3) Enhanced the distinction and boundary clarity between lineages. By employing fine-grained modification strategies (such as breakpoint reassignment), this embodiment effectively optimizes inter-class separation. The clustering effect was evaluated using a simulated Davies-Bouldin index (SDBI) (a lower index indicates more compact intra-class clusters and better inter-class separation). Results showed that for SARS-CoV-2, the average SDBI value in this embodiment was 6.15, significantly better than Pangolin's 7.71; for monkeypox virus, the average SDBI value was 1.77, better than Nextclade's 1.84. This indicates that this embodiment can more clearly define the biological boundaries between different evolutionary lineages, reducing the lineage overlap phenomenon commonly found in existing technologies.
[0158] Based on the same technical concept, embodiments of this application also provide a device for classifying pathogen evolutionary lineages, such as... Figure 5 As shown, the device includes: Acquisition module 501 is used to acquire multiple genomic sequences of the pathogen; The first construction module 502 is used to construct a weighted haplotype network based on the mutation characteristics of the genome sequence; wherein, the nodes in the weighted haplotype network are haplotypes, and each haplotype represents a group of genome sequences with the same mutation characteristics; the weight of the edges between nodes is calculated based on the genetic similarity between the corresponding haplotypes; The first partitioning module 503 is used to perform community structure detection on the weighted monomer network based on the modularity optimization algorithm, and to divide the monomers in the weighted monomer network into multiple coarse-grained clusters to obtain coarse-grained partitioning results. The second division module 504 is used to refine the coarse-grained division result to obtain the final pathogen evolutionary lineage division result. Specifically, the second partitioning module is used to perform the fine-grained modification in the following manner: In the weighted monomer network, monomers that are connected to only one downstream monomer and are classified into different coarse-grained clusters with that downstream monomer are reassigned to the cluster with the highest node-cluster similarity. And / or, For coarse-grained clusters where the number of genomic sequences corresponding to haplotypes is less than a preset threshold, the mutation similarity between the coarse-grained cluster and its parent cluster is calculated. If the mutation similarity is greater than a preset merging threshold, the coarse-grained cluster is merged into its parent cluster.
[0159] Optionally, when the first construction module 502 is used to construct a weighted haplotype network based on the mutation characteristics of the genome sequence, it is specifically used for: For each pathogen's genome sequence, the mutations of that genome sequence relative to the reference sequence are extracted to form a mutation set consisting of mutation sites and the data after mutation. Pathogens with the same mutation set are grouped into the same haplotype, resulting in multiple haplotypes. Based on the mutation sets corresponding to each of the complete genome sequences of the pathogen, a haplotype network is constructed using a haplotype network building algorithm; For any two haplotypes that are connected in the haplotype network, the Jaccard coefficient is calculated as the genetic similarity between the two haplotypes based on the mutation sets of the two haplotypes. Based on the connection relationship and genetic similarity between the two haplotypes, weights are calculated and assigned to the edges between the two haplotypes in the haplotype network to obtain a weighted haplotype network; wherein, the genetic similarity between haplotypes is positively correlated with the weights.
[0160] Optionally, when the second partitioning module 504 performs community structure detection on the weighted monomer network based on the modularity optimization algorithm, and partitions the monomers in the weighted monomer network into multiple coarse-grained clusters to obtain coarse-grained partitioning result clusters, it is specifically used for: Each monomer in the weighted monomer network is initialized as an initial cluster center for an independent cluster; An iterative process is performed, which includes a local movement phase and a network compression phase. During the local movement phase, for each monomer in the current monomer network, the modularity gain caused by moving to any adjacent cluster is calculated. If there is a positive modularity gain, it is moved to the adjacent cluster that brings the maximum modularity gain; if there is no movement that brings a positive modularity gain, its current cluster affiliation remains unchanged. During the network compression phase, all monomers belonging to the same cluster after local movement are compressed into a new supernode, and a new compressed network is constructed based on the original weight relationship. The compressed network is used as a new weighted monolithic network, and the local movement phase and the network compression phase are repeated. The iteration stops when a positive modularity gain cannot be obtained through node movement in the weighted monolithic network. The cluster partitioning result obtained after the final iteration is used as the coarse-grained partitioning result.
[0161] Optionally, the second partitioning module 504, when reassigning a monomer in the weighted monomer network that is connected to only one downstream monomer and is classified in a different coarse-grained cluster with that downstream monomer to the cluster with the highest node-cluster similarity, specifically performs the following: From each of the coarse-grained clusters, a monomer type that meets the following condition is identified as a node to be reallocated: in the weighted monomer type network, the monomer type has one and only one directly connected downstream monomer type, and the monomer type and the downstream monomer type are located in different coarse-grained clusters. For each node to be reassigned, a first Jaccard coefficient is calculated between the node to be reassigned and its current first coarse-grained cluster; wherein, the first Jaccard coefficient is calculated based on the mutation set of the node to be reassigned and the mutation set shared by all haplotypes in the first coarse-grained cluster; the first Jaccard coefficient is used to represent the node-cluster similarity between the node to be reassigned and its current first coarse-grained cluster. Calculate the second Jaccard coefficient between the node to be reassigned and the second coarse-grained cluster to which its downstream haplotype belongs; wherein, the second Jaccard coefficient is calculated based on the mutation set of the node to be reassigned and the mutation set common to all haplotypes in the second coarse-grained cluster; the second Jaccard coefficient is used to represent the node-cluster similarity between the node to be reassigned and the second coarse-grained cluster to which its downstream haplotype belongs; Compare the first Jaccard coefficient with the second Jaccard coefficient; If the second Jaccard coefficient is greater than the first Jaccard coefficient, the node to be reassigned is reassigned from the first coarse-grained cluster to the second coarse-grained cluster; if the second Jaccard coefficient is less than or equal to the first Jaccard coefficient, the node's affiliation in the first coarse-grained cluster remains unchanged.
[0162] Optionally, the second partitioning module 504, when calculating the mutation similarity between a coarse-grained cluster and its parent cluster for a coarse-grained cluster containing a haplotype whose number of corresponding genomic sequences is less than a preset threshold, and merging the coarse-grained cluster into its parent cluster if the mutation similarity is greater than a preset merging threshold, specifically uses the following: From each coarse-grained cluster, identify small-scale clusters whose number of genomic sequences corresponding to haplotypes is less than a preset threshold; For each of the small-scale clusters, a coarse-grained cluster with which it has a connection is identified in the weighted monomorphic network as its parent cluster; Determine the common mutation set of the small-scale clusters; the common mutation set consists of mutation sites and mutated data shared by monomers of not less than a preset proportion within the cluster; Determine the common mutation set of the parent cluster; Based on the common mutation set of the small-scale cluster and the common mutation set of the parent cluster, the cluster-cluster Jaccard coefficient is calculated as the mutation similarity between the small-scale cluster and the parent cluster; Compare the mutation similarity with the preset merging threshold; If the mutation similarity is greater than the merging threshold, the small cluster is merged into the parent cluster; if the mutation similarity is less than or equal to the merging threshold, the small cluster remains independent.
[0163] Optionally, the device further includes: The generation module is used to take the genome sequence of the newly added weighted haplotype network as incremental data, update the structure of the weighted haplotype network based on the newly added genome sequence, re-perform the community structure detection and fine-grained modification on the updated weighted haplotype network, and generate a candidate new lineage set independent of the current lineage set; the current lineage set is the pathogen evolutionary lineage division result before the addition of incremental data; the candidate new lineage set consists of multiple candidate new evolutionary lineages; The calculation module is used to calculate the aggregation ratio and dispersion ratio from the current evolutionary lineage to the candidate new evolutionary lineage for any given current evolutionary lineage and a candidate new evolutionary lineage; wherein, the aggregation ratio is the proportion of the number of genomic sequences corresponding to haplotypes shared by the current evolutionary lineage and the candidate new evolutionary lineage to the total number of genomic sequences corresponding to haplotypes contained in the candidate new evolutionary lineage; the dispersion ratio is the proportion of the number of genomic sequences corresponding to haplotypes shared by the current evolutionary lineage and the candidate new evolutionary lineage to the total number of genomic sequences corresponding to haplotypes contained in the current evolutionary lineage; The determination module is used to perform hierarchical lineage determination on the genomic sequences corresponding to haplotypes in the candidate new lineage set based on the aggregation ratio and the dispersion ratio, so as to correct the haplotype attribution in the candidate new lineage set and generate an updated pathogen evolution lineage classification result. The step of performing hierarchical lineage classification determination on haplotypes in the candidate new lineage set includes: If a monomer type is assigned to the root lineage in the current lineage set and is classified into a non-root lineage in the candidate new lineage set, then the assignment change is accepted when the aggregation ratio from the root lineage to the non-root lineage is greater than or equal to the first preset threshold; otherwise, the monomer type is forced to remain in the root lineage after the update. If a haplotype is assigned to any first-generation lineage in the current lineage set and is assigned to the root lineage in the candidate new lineage set, then if the dispersion ratio from the first-generation lineage to the root lineage is less than a second preset threshold, the assignment change is accepted; otherwise, the haplotype is forced to remain assigned to the original first-generation lineage after the update. For a single type that does not belong to the root lineage in the current lineage set and whose original lineage and root lineage dispersion ratio fails to meet the second preset threshold, its affiliation in the candidate new lineage set is modified to the candidate new evolutionary lineage that has the largest aggregation ratio with its original lineage in the current lineage set.
[0164] Optionally, after generating the updated pathogen evolutionary lineage classification results, the updated pathogen evolutionary lineage classification results are recorded as the newly formed lineage set, and the candidate new evolutionary lineages in the updated pathogen evolutionary lineage classification results are recorded as newly formed lineages; the device further includes: The second construction module is used to construct a cross-relationship matrix between the newly formed lineage and the current evolutionary lineage; wherein, the rows of the cross-relationship matrix correspond to the current evolutionary lineage, the columns correspond to the newly formed lineage, and the element values in the cross-relationship matrix are the number of genomic sequences corresponding to haplotypes shared by the lineages in the corresponding rows and columns; The inheritance module is used to traverse the cross-relation matrix. For each new lineage in the new lineage set, if the new lineage is a first-generation lineage directly derived from the root lineage, the current evolutionary lineage with the largest proportion of genomic sequence sources corresponding to its haplotype is determined according to the cross-relation matrix, and the name identifier is inherited from the current evolutionary lineage. If the name identifier is not occupied, it is allocated; otherwise, a new first-generation lineage name identifier is allocated to it. The search module is used to, if the newly formed lineage does not belong to the first-generation lineage, search in the cross-relationship matrix for whether there exists a current evolutionary lineage such that the number of genomic sequences corresponding to haplotypes shared by the newly formed lineage and the current evolutionary lineage is simultaneously the maximum value of the row where the current evolutionary lineage is located and the maximum value of the column where the newly formed lineage is located; if such a lineage exists, the newly formed lineage is determined to be the direct successor of the current evolutionary lineage, and a name identifier of the current evolutionary lineage is assigned to it; The allocation module is used to allocate a new incremental name identifier to the new lineage if it is neither a first-generation lineage directly derived from the root lineage nor a direct inheritor of any current evolutionary lineage. The new incremental name identifier is generated by keeping the parent identifier of the first-generation lineage unchanged and setting the offspring sequence number in the first-generation lineage to the increment of the current maximum sequence number. The association module is used to associate the final determined name identifiers with the set of new lineages after all name identifiers of all new lineages have been assigned, so as to output a set of new lineages that has been associated with the name identifiers.
[0165] Figure 6 A schematic diagram of an electronic device provided in this application embodiment includes: a processor 601, a memory 602, and a bus 606. The memory 602 stores machine-readable instructions executable by the processor 601. When the electronic device runs the above-described information processing method, the processor 601 and the memory 602 communicate through the bus 606. The processor 601 executes the machine-readable instructions to perform the steps of the method described in Embodiment 1.
[0166] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps described in Embodiment 1.
[0167] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, electronic devices, and computer-readable storage media described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0168] In the several embodiments provided in this application, it should be understood that the disclosed methods, apparatuses, electronic devices, and computer-readable storage media can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or modules may be electrical, mechanical, or other forms.
[0169] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0170] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0171] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0172] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of protection of the claims.
Claims
1. A method for classifying the evolutionary lineage of pathogens, characterized in that, include: Obtain multiple genome sequences of the pathogen; Based on the mutation characteristics of the genome sequences, a weighted haplotype network is constructed; wherein, the nodes in the weighted haplotype network are haplotypes, and each haplotype represents a group of genome sequences with the same mutation characteristics; the weights of the edges between nodes are calculated based on the genetic similarity between the corresponding haplotypes. The community structure of the weighted monomer network is detected based on the modularity optimization algorithm, and the monomers in the weighted monomer network are divided into multiple coarse-grained clusters to obtain the coarse-grained partitioning result. The coarse-grained classification results are then refined to obtain the final pathogen evolutionary lineage classification results; wherein, the fine-grained refinement includes: In the weighted monomer network, monomers that are connected to only one downstream monomer and are classified into different coarse-grained clusters with that downstream monomer are reassigned to the cluster with the highest node-cluster similarity. And / or, For coarse-grained clusters where the number of genomic sequences corresponding to haplotypes is less than a preset threshold, the mutation similarity between the coarse-grained cluster and its parent cluster is calculated. If the mutation similarity is greater than a preset merging threshold, the coarse-grained cluster is merged into its parent cluster.
2. The method according to claim 1, characterized in that, The construction of a weighted haplotype network based on the mutation characteristics of the genome sequence includes: For each pathogen's genome sequence, the mutations of that genome sequence relative to the reference sequence are extracted to form a mutation set consisting of mutation sites and the data after mutation. Pathogens with the same mutation set are grouped into the same haplotype, resulting in multiple haplotypes. Based on the mutation sets corresponding to each of the complete genome sequences of the pathogen, a haplotype network is constructed using a haplotype network building algorithm; For any two haplotypes that are connected in the haplotype network, the Jaccard coefficient is calculated as the genetic similarity between the two haplotypes based on the mutation sets of the two haplotypes. Based on the connection relationship and genetic similarity between the two haplotypes, weights are calculated and assigned to the edges between the two haplotypes in the haplotype network to obtain a weighted haplotype network; wherein, the genetic similarity between haplotypes is positively correlated with the weights.
3. The method according to claim 1, characterized in that, The community structure detection of the weighted monomer network based on the modularity optimization algorithm divides the monomers in the weighted monomer network into multiple coarse-grained clusters, obtaining coarse-grained partitioning results, including: Each monomer in the weighted monomer network is initialized as an initial cluster center for an independent cluster; An iterative process is performed, which includes a local movement phase and a network compression phase. During the local movement phase, for each monomer in the current monomer network, the modularity gain caused by moving to any adjacent cluster is calculated. If there is a positive modularity gain, it is moved to the adjacent cluster that brings the maximum modularity gain; if there is no movement that brings a positive modularity gain, its current cluster affiliation remains unchanged. During the network compression phase, all monomers belonging to the same cluster after local movement are compressed into a new supernode, and a new compressed network is constructed based on the original weight relationship. The compressed network is used as a new weighted monolithic network, and the local movement phase and the network compression phase are repeated. The iteration stops when a positive modularity gain cannot be obtained through node movement in the weighted monolithic network. The cluster partitioning result obtained after the final iteration is used as the coarse-grained partitioning result.
4. The method according to claim 2, characterized in that, The step of reassigning a monomer in the weighted monomer network that is connected to only one downstream monomer and is classified in a different coarse-grained cluster to the cluster with the highest node-cluster similarity includes: From each of the coarse-grained clusters, a monomer type that meets the following condition is identified as a node to be reallocated: in the weighted monomer type network, the monomer type has one and only one directly connected downstream monomer type, and the monomer type and the downstream monomer type are located in different coarse-grained clusters. For each node to be reassigned, a first Jaccard coefficient is calculated between the node to be reassigned and its current first coarse-grained cluster; wherein, the first Jaccard coefficient is calculated based on the mutation set of the node to be reassigned and the mutation set shared by all haplotypes in the first coarse-grained cluster; the first Jaccard coefficient is used to represent the node-cluster similarity between the node to be reassigned and its current first coarse-grained cluster. Calculate the second Jaccard coefficient between the node to be reassigned and the second coarse-grained cluster to which its downstream haplotype belongs; wherein, the second Jaccard coefficient is calculated based on the mutation set of the node to be reassigned and the mutation set common to all haplotypes in the second coarse-grained cluster; the second Jaccard coefficient is used to represent the node-cluster similarity between the node to be reassigned and the second coarse-grained cluster to which its downstream haplotype belongs; Compare the first Jaccard coefficient with the second Jaccard coefficient; If the second Jaccard coefficient is greater than the first Jaccard coefficient, the node to be reassigned is reassigned from the first coarse-grained cluster to the second coarse-grained cluster; if the second Jaccard coefficient is less than or equal to the first Jaccard coefficient, the node's affiliation in the first coarse-grained cluster remains unchanged.
5. The method according to claim 1, characterized in that, For coarse-grained clusters containing haplotypes with fewer than a preset threshold number of genomic sequences, the mutation similarity between the cluster and its parent cluster is calculated. If the mutation similarity is greater than a preset merging threshold, the coarse-grained cluster is merged into its parent cluster, including: From each coarse-grained cluster, identify small-scale clusters whose number of genomic sequences corresponding to haplotypes is less than a preset threshold; For each of the small-scale clusters, a coarse-grained cluster with which it has a connection relationship is determined in the weighted monomorphic network as its parent cluster; Determine the common mutation set of the small-scale cluster; the common mutation set consists of mutation sites and data after mutation shared by monomers of not less than a preset proportion within the cluster; Determine the common mutation set of the parent cluster; Based on the common mutation set of the small-scale cluster and the common mutation set of the parent cluster, the cluster-cluster Jaccard coefficient is calculated as the mutation similarity between the small-scale cluster and the parent cluster; Compare the mutation similarity with the preset merging threshold; If the mutation similarity is greater than the merging threshold, the small cluster is merged into the parent cluster; if the mutation similarity is less than or equal to the merging threshold, the small cluster remains independent.
6. The method according to claim 1, characterized in that, After refining the coarse-grained classification results to obtain the final pathogen evolutionary lineage classification results, the method further includes: The genomic sequences newly added to the weighted haplotype network are used as incremental data. The structure of the weighted haplotype network is updated based on these newly added genomic sequences. The community structure detection and fine-grained modification are then re-executed on the updated weighted haplotype network to generate a candidate new lineage set independent of the current lineage set. The current lineage set is the pathogen evolutionary lineage division result before the addition of incremental data. The candidate new lineage set consists of multiple candidate new evolutionary lineages. For any current evolutionary lineage and a candidate new evolutionary lineage, calculate the aggregation ratio and dispersion ratio from the current evolutionary lineage to the candidate new evolutionary lineage; wherein, the aggregation ratio is the proportion of the number of genomic sequences corresponding to haplotypes shared by the current evolutionary lineage and the candidate new evolutionary lineage to the total number of genomic sequences corresponding to haplotypes contained in the candidate new evolutionary lineage; the dispersion ratio is the proportion of the number of genomic sequences corresponding to haplotypes shared by the current evolutionary lineage and the candidate new evolutionary lineage to the total number of genomic sequences corresponding to haplotypes contained in the current evolutionary lineage; Based on the aggregation ratio and the dispersion ratio, the genomic sequences corresponding to haplotypes in the candidate new lineage set are subjected to hierarchical lineage classification to correct the classification of haplotypes in the candidate new lineage set, thereby generating updated pathogen evolutionary lineage classification results. The step of performing hierarchical lineage classification determination on haplotypes in the candidate new lineage set includes: If a monomer type is assigned to the root lineage in the current lineage set and is classified into a non-root lineage in the candidate new lineage set, then the assignment change is accepted when the aggregation ratio from the root lineage to the non-root lineage is greater than or equal to the first preset threshold; otherwise, the monomer type is forced to remain in the root lineage after the update. If a haplotype is assigned to any first-generation lineage in the current lineage set and is assigned to the root lineage in the candidate new lineage set, then if the dispersion ratio from the first-generation lineage to the root lineage is less than a second preset threshold, the assignment change is accepted; otherwise, the haplotype is forced to remain assigned to the original first-generation lineage after the update. For a single type that does not belong to the root lineage in the current lineage set and whose original lineage and root lineage dispersion ratio fails to meet the second preset threshold, its affiliation in the candidate new lineage set is modified to the candidate new evolutionary lineage that has the largest aggregation ratio with its original lineage in the current lineage set.
7. The method according to claim 6, characterized in that, After generating the updated pathogen evolutionary lineage classification results, the updated pathogen evolutionary lineage classification results are recorded as the newly formed lineage set, and the candidate new evolutionary lineages in the updated pathogen evolutionary lineage classification results are recorded as newly formed lineages; the method further includes: Construct a cross-relationship matrix between the newly formed lineage and the current evolutionary lineage; wherein, the rows of the cross-relationship matrix correspond to the current evolutionary lineage, the columns correspond to the newly formed lineage, and the element values in the cross-relationship matrix are the number of genomic sequences corresponding to haplotypes shared by the lineages in the corresponding rows and columns; Traverse the cross-relation matrix. For each new lineage in the new lineage set, if the new lineage is a first-generation lineage directly derived from the root lineage, determine the current evolutionary lineage with the largest proportion of genomic sequence source corresponding to its haplotype according to the cross-relation matrix, and inherit the name identifier from the current evolutionary lineage; if the name identifier is not occupied, it is allocated; otherwise, a new first-generation lineage name identifier is allocated to it. If the newly formed lineage does not belong to the first generation lineage, then in the cross-relationship matrix, it is searched to see if there exists a current evolutionary lineage such that the number of genomic sequences corresponding to the haplotypes shared by the newly formed lineage and the current evolutionary lineage is simultaneously the maximum value of the row where the current evolutionary lineage is located and the maximum value of the column where the newly formed lineage is located; if such a lineage exists, then the newly formed lineage is determined to be the direct successor of the current evolutionary lineage, and a name identifier of the current evolutionary lineage is assigned to it; If the new lineage is neither a first-generation lineage directly derived from the root lineage nor can it be determined as a direct inheritor of any current evolutionary lineage, then a new incremental name identifier is assigned to it; the new incremental name identifier is generated by keeping the parent identifier of its first-generation lineage unchanged and setting the offspring sequence number within the first-generation lineage to the increment of the current maximum sequence number. After all new lineage names have been assigned, the final assigned name identifiers are associated with the new lineage set to output a new lineage set associated with the name identifiers.
8. A device for classifying the evolutionary lineage of pathogens, characterized in that, include: The acquisition module is used to acquire multiple genomic sequences of the pathogen; The first construction module is used to construct a weighted haplotype network based on the mutation characteristics of the genome sequence; wherein, the nodes in the weighted haplotype network are haplotypes, and each haplotype represents a group of genome sequences with the same mutation characteristics; the weight of the edges between nodes is calculated based on the genetic similarity between the corresponding haplotypes; The first partitioning module is used to perform community structure detection on the weighted monomer network based on the modularity optimization algorithm, and to divide the monomers in the weighted monomer network into multiple coarse-grained clusters to obtain coarse-grained partitioning results. The second partitioning module is used to refine the coarse-grained partitioning results to obtain the final pathogen evolutionary lineage partitioning results. Specifically, the second partitioning module is used to perform the fine-grained modification in the following manner: In the weighted monomer network, monomers that are connected to only one downstream monomer and are classified into different coarse-grained clusters with that downstream monomer are reassigned to the cluster with the highest node-cluster similarity. And / or, For coarse-grained clusters where the number of genomic sequences corresponding to haplotypes is less than a preset threshold, the mutation similarity between the coarse-grained cluster and its parent cluster is calculated. If the mutation similarity is greater than a preset merging threshold, the coarse-grained cluster is merged into its parent cluster.
9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is in operation, the processor communicates with the memory via the bus, and the machine-readable instructions, when executed by the processor, perform the steps of the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the method as described in any one of claims 1 to 7.