Efficient heterogeneous information network graph structured clustering method and system based on meta-path enhancement
By building a meta-path transformation map in a heterogeneous information network and combining independent and non-independent clustering models, the application difficulties of traditional SCAN methods in heterogeneous networks are solved, efficient heterogeneous network clustering and hub point recognition are achieved, and rapid processing of large-scale networks is supported.
Patent Information
- Application Number
- CN202510351470.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional SCAN method assumes that the input graph is an isomorphic network, difficult to directly apply to heterogeneous information networks, and it is difficult to distinguish between hub points and free points in heterogeneous information networks.
By combining metapathic constraints with structural similarity calculations, two heterogeneous network structure clustering models are constructed, two independent and non-independent, two-way metapath instance search are used to build metapathic transformation maps, and redundant calculations are eliminated through boundary pruning strategies to identify hub nodes and outliers.
It significantly improves the accuracy of heterogeneous network clustering, supports minute-level processing of billion-scale networks, improves hub recognition accuracy, and provides multi-dimensional node role evaluation.
Smart Images

Figure CN120277252A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer science and technology, and particularly to an efficient heterogeneous information network graph structured clustering method and system based on meta-path enhancement. Background Art
[0002] With the development of technologies such as social networks, mobile Internet, and Internet of Things, the scale and speed of data acquisition have exploded, and technologies related to big data have become a hot topic in the world today. In big data analysis, as an abstract data structure that effectively describes the correlation between data, graphs play an increasingly important role and are widely used in various fields such as social networks, e-commerce, urban transportation, communication networks, biochemistry, etc. The graph clustering problem, as one of the basic problems in graph data analysis and processing, has received extremely high attention from foreign researchers. Given a graph, the ultimate goal of a graph clustering algorithm is to group the vertices in the graph so that there are dense edges between the vertices in the same group, while there are few edges between the vertices belonging to different groups. Graph clustering has many applications because it can discover the characteristics of the closely connected clustering structure in graph data.
[0003] Current researchers have proposed many different graph clustering methods. They include modularity-based methods, graph partitioning, and density-based methods. However, most of these clustering methods only focus on the calculation of clustering and do not distinguish the specific roles of vertices that do not belong to any cluster. Some of these vertices are hub points that connect many clusters but do not belong to any cluster, and the other part is free points that have only weak associations with specific clusters. Distinguishing hub points and free points is important for mining various complex graph data. To solve this problem, researchers have proposed a graph structure clustering method (SCAN method). SCAN defines the structural similarity between adjacent vertices. If the structural similarity between two vertices is not less than a given parameter ε (similarity threshold), they are considered to be associated, represented by their similarity. To construct clusters, SCAN detects a special type of vertex called a core. A core is a point that is similar to many vertices and is regarded as the key to a cluster. The lowest threshold of the number of similar neighbors for a vertex to become a core point is determined by a given parameter μ. A cluster is constructed by a core continuously expanding according to similar neighbors. And the vertices that are not in any cluster are further divided into hub points and outliers.
[0004] Currently, in many practical applications, such as literature networks, social networks, and knowledge graphs, data usually exists in the form of heterogeneous information networks. These networks contain multiple types of entities (vertices) and different types of relationships (edges) between them. The traditional SCAN method (a graph structure clustering method) assumes that the input graph is a homogeneous network, that is, all vertices are of the same type, which limits its application in heterogeneous information networks. In addition, the complexity of heterogeneous information networks also lies in that even vertices of the same type may not have direct connections, let alone have common neighbors and more complex definitions based on this. These characteristics make it difficult to directly apply the traditional SCAN method to heterogeneous information networks. Summary of the Invention
[0005] The present invention provides an efficient heterogeneous information network graph structure clustering method and system based on meta-path enhancement. By combining meta-path semantic constraints with structural similarity calculation and constructing two heterogeneous network structure clustering models, independent and non-independent, the accuracy of heterogeneous network clustering is significantly improved. While maintaining the hub recognition ability of the SCAN algorithm, it supports minute-level processing of networks on the scale of billions, and can be widely applied to the fields of social network analysis, knowledge graph mining, and recommendation system optimization.
[0006] The present invention provides an efficient heterogeneous information network graph structure clustering method based on meta-path enhancement, including:
[0007] S1. Obtain and parse multi-type nodes in the heterogeneous network, define a vertex type set and an edge type set to generate a candidate meta-path set, and select a target meta-path based on the target clustering vertices;
[0008] S2. Construct a meta-path conversion graph by using the method of bidirectional meta-path instance search according to the target meta-path;
[0009] S3. Implement non-independent and independent dual-mode clustering models on the meta-path conversion graph to obtain the core clustering result; among them, for the independent clustering model, use independent path verification based on recursive segmentation to ensure the complete independence between path instances, and eliminate redundant calculations through the boundary pruning strategy;
[0010] S4. Identify hub nodes and outliers based on the obtained core clustering result. The hub nodes need to meet the cross-cluster adjacency condition, and the outliers need to be excluded from the core extension path. Construct a multi-dimensional node role feature vector for visual annotation to complete the structure clustering of the heterogeneous information network.
[0011] Further, the step S1 specifically includes:
[0012] S101. Load the heterogeneous information network from the data source. The heterogeneous information network contains various types of nodes and edges. By parsing the network structure, extract each type of node and its attribute information;
[0013] S102. According to the node types and edge types in the heterogeneous information network, define a vertex type set and an edge type set. The vertex type set contains all different types of nodes in the network, and the edge type set contains all different types of edges. Based on the vertex type set and the edge type set, generate all possible meta-path sets by combining different types of vertices and edges. Among them, a meta-path is a sequence of vertex types connected in a specific order, used to represent the semantic relationship between nodes;
[0014] S103. Determine the target clustering vertex type, that is, the node type that needs to perform clustering analysis. Screen out the meta-paths in the generated meta-path set whose starting vertex type and ending vertex type are the same as the target clustering vertex type. Generate a candidate meta-path set through network pattern traversal, and screen out the high-frequency meta-paths. Finally, select the meta-path that maximizes the average connection density between target type nodes as the target meta-path for subsequent clustering analysis tasks.
[0015] Further, the step S2 specifically includes:
[0016] S201. Determine the node type with the least number of vertices in the target meta-path as the splitting point type;
[0017] S202. Perform depth-first search along the forward and reverse directions of the target meta-path respectively, and collect the reachable target type node sets. Among them, define the forward direction as the sequential direction from the splitting point type of the target meta-path to the ending vertex type, and the reverse direction is the reverse order direction from the splitting point type to the starting vertex type;
[0018] S203. Perform a Cartesian product connection on the forward target type node set and the reverse target type node set to generate a complete meta-path conversion graph.
[0019] Further, the step S3 specifically includes:
[0020] S301. Implement a non-independent clustering model on the meta-path conversion graph to obtain the first core clustering result;
[0021] S302. Implement an independent clustering model on the meta-path conversion graph to obtain the second core clustering result.
[0022] Further, the step S301 specifically includes:
[0023] S3011. Define the structural similarity between nodes in the non-independent mode. For any two nodes u and v in the meta-path conversion graph, their structural similarity σ P (u, v) is defined as the ratio of the number of their common neighbors to the geometric mean of the number of their neighbors, and the calculation formula is: where σ P (u, v) represents the similarity between nodes u and v, and N P (u) represents the set of neighbor nodes of node u based on the meta-path P;
[0024] S3012. Identify core nodes by setting a similarity threshold ∈ and a minimum number of neighbors μ. If the ∈-neighbor number of a node u is greater than or equal to μ, then u is defined as a core node; where the ∈-neighbor number of the node u is the number of neighbors whose structural similarity with u is greater than or equal to ∈;
[0025] S3013. Continuously expand and cluster through the core nodes until no further expansion is possible, and obtain the first core clustering result.
[0026] Furthermore, in the step S302, the independent clustering model process adopts the overall framework of the non-independent model and makes the following improvements:
[0027] The independent mode clustering model introduces an independent path verification mechanism based on recursive segmentation to ensure that the connection paths between nodes do not share intermediate entities; the independent path verification mechanism first recursively divides the target meta-path into a left half-path P l and a right half-path P r , then for each candidate common node w, calculate the set of left half-path instances from u to w and the set of right half-path instances from v to w respectively, then construct a candidate node matching matrix, and enumerate the path combinations that do not share intermediate nodes through the backtracking method. Finally, when there are completely independent (u, w), (v, w), and (u, v) three groups of path instances, determine that w is an independent common neighbor;
[0028] The calculation of the structural similarity of the independent mode clustering model introduces an independence constraint: where represents the independent path similarity between nodes u and v, and I P (u, v) represents the set of common neighbor nodes connected by independent meta-path instances between u and v;
[0029] During the clustering process of the independent mode clustering model, the boundary pruning strategy is implemented through boundary maintenance to eliminate redundant calculations. The process is as follows: pre-compute the core candidate set in the non-independent mode as the upper bound; maintain the upper and lower bounds (lb, ub) of the number of similar neighbors of each node; directly prune non-core nodes when ub < μ; implement a delayed verification strategy for boundary nodes, only when lb +Perform a complete independence check when the number of unvalidated neighbors ≥ μ.
[0030] Further, the step S4 specifically includes:
[0031] S401. Identify hub nodes and outliers based on the obtained first core clustering result and second core clustering result;
[0032] Identify hub nodes: After completing the construction of the core clustering, traverse all the nodes that have not been assigned to the core clustering result, and then check whether these nodes are connected to the nodes in multiple core clustering results. If so, mark this node as a hub node;
[0033] Identify outliers: After completing the construction of the core clustering, traverse all the nodes that have not been assigned to the core clustering result, and then check whether these nodes are connected to the nodes in multiple core clustering results. If a node neither meets the conditions of a hub node nor belongs to any core clustering result, mark it as an outlier;
[0034] S402. Construct a multi-dimensional node role feature vector to achieve visual annotation and complete the structure clustering of the heterogeneous information network.
[0035] The present invention also provides an efficient heterogeneous information network graph structure clustering system based on meta-path enhancement, including:
[0036] An acquisition module, configured to acquire and parse multiple types of nodes in the heterogeneous network, define a vertex type set and an edge type set to generate a candidate meta-path set, and select a target meta-path based on the target clustering vertex;
[0037] A construction module, configured to construct a meta-path conversion graph by means of bidirectional meta-path instance search according to the target meta-path;
[0038] A dual-mode clustering module, configured to implement a non-independent and independent dual-mode clustering model on the meta-path conversion graph to obtain a core clustering result; wherein, for the independent clustering model, an independent path verification based on recursive segmentation is adopted to ensure the complete independence between path instances, and redundant calculations are eliminated through a boundary pruning strategy;
[0039] An identification and annotation module, applied to identify hub nodes and outliers based on the obtained core clustering result, where the hub nodes need to meet the cross-cluster adjacency condition, the outliers need to be excluded from the core extension path, construct a multi-dimensional node role feature vector to achieve visual annotation, and complete the structure clustering of the heterogeneous information network.
[0040] The present invention also provides a computer device, including a memory and a processor, where the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0041] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0042] The beneficial effects of the present invention are as follows:
[0043] The present invention obtains and analyzes multi-type nodes in a heterogeneous network, defines a vertex type set and an edge type set to generate a candidate meta-path set, and selects a target meta-path based on target clustering vertices; constructs a meta-path conversion graph by adopting a two-way meta-path instance search method according to the target meta-path; implements a non-independent and independent dual-mode clustering model on the meta-path conversion graph to obtain a core clustering result; finally, based on the obtained core clustering result, hub nodes and outliers are identified, and a multi-dimensional node role feature vector is constructed to realize visual annotation, thereby completing the structural clustering of the heterogeneous information network. In the present invention,
[0044] (1) A two-way search conversion graph construction method is created, and compared with the traditional DFS, the construction efficiency is increased by 5.7 times, and the memory occupancy is reduced by 63%.
[0045] (2) The F1 value of the dual-mode clustering model in the DBLP network reaches 0.91, which is increased by 29% compared with the single mode.
[0046] (3) The dynamic pruning strategy reduces 85% of the redundant calculations and supports the processing of a heterogeneous network with billions of edges in minutes.
[0047] (4) The independent path verification framework accurately eliminates 92% of the falsely high similarity relationships and improves the hub identification accuracy.
[0048] (5) The multi-dimensional role annotation system can quantitatively evaluate the node influence and provides a new dimension for community discovery. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention.
[0050] Figure 2 It is a simple example diagram of a heterogeneous information network in the present invention.
[0051] Figure 3 It is a schematic diagram of the isomorphic graph constructed by the two-way search conversion process in the present invention.
[0052] Figure 4 It is a schematic diagram of the two-way search process and the pivot node in the present invention.
[0053] Figure 5 It is a schematic diagram of the recursive independent path verification process in the present invention.
[0054] Figure 6Schematic diagram of the clustering result in the present invention.
[0055] Figure 7 Schematic diagram of the device structure according to an embodiment of the present invention.
[0056] Figure 8 Schematic diagram of the internal structure of a computer device according to an embodiment of the present invention.
[0057] The realization, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0058] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0059] The present invention is used for efficient clustering analysis of heterogeneous information networks. First, a meta-path pattern is defined according to the target clustering vertex type, a meta-path transformation graph is constructed through a search-join transformation paradigm, and the heterogeneous network is mapped into a homogeneous network structure. For the transformed network, a two-stage clustering model is proposed: in the non-independent mode, core nodes are identified through the structural similarity of meta-path neighbors, and clusters are formed by connectivity expansion; in the independent mode, a path independence verification mechanism is introduced to ensure that the connection paths between nodes do not share intermediate nodes, effectively avoiding the single-point failure problem. Based on the definition of core nodes, the system adopts a two-layer boundary maintenance strategy, realizes computational pruning by dynamically updating the upper and lower bounds of neighbor similarity, and designs a parallel verification framework to accelerate the discrimination process of independent paths. By combining meta-path semantic constraints with structural similarity calculation and constructing two heterogeneous network structure clustering models of independent and non-independent types, the present invention significantly improves the accuracy of heterogeneous network clustering, supports minute-level processing of networks on the scale of billions while maintaining the hub identification ability of the SCAN algorithm, and can be widely applied to the fields of social network analysis, knowledge graph mining and recommendation system optimization.
[0060] As Figure 1 shown, the present invention provides an efficient heterogeneous information network graph-structured clustering method based on meta-path enhancement, including:
[0061] S1. Obtain and parse multi-type nodes in the heterogeneous network, define a vertex type set and an edge type set to generate a candidate meta-path set, and select a target meta-path based on the target clustering vertex.
[0062] Specifically, it includes the following steps:
[0063] S101. Load the heterogeneous information network from the data source. The heterogeneous information network contains multiple types of nodes and edges. By parsing the network structure, each type of node and its attribute information are extracted.
[0064] S102. Define a vertex type set and an edge type set according to the node types and edge types in the heterogeneous information network. The vertex type set contains all different types of nodes in the network, and the edge type set contains all different types of edges. Based on the vertex type set and the edge type set, generate all possible meta-path sets by combining different types of vertices and edges. Among them, a meta-path is a sequence of vertex types connected in a specific order, which is used to represent the semantic relationship between nodes.
[0065] S103. Determine the target clustering vertex type, that is, the node type that needs to perform clustering analysis. Filter out the meta-paths from the generated meta-path set whose starting vertex type and ending vertex type are the same as the target clustering vertex type. Generate a candidate meta-path set by traversing the network pattern, and filter out the high-frequency meta-paths. Finally, select the meta-path that maximizes the average connection density between nodes of the target type as the target meta-path for subsequent clustering analysis tasks.
[0066] As described in the above steps S101 - S103, refer to Figure 2 , parse the network pattern, define the vertex type set A and the edge type set R, generate the candidate meta-path set P. The candidate meta-path set P is a symmetric meta-path candidate set with a length ≤ 4. Count the number of instances of each meta-path, and retain the top 20% high-frequency paths. Then select the optimal meta-path P based on the target clustering vertex type, satisfying P[0] = P[|P| - 1] and having the maximum connection density. Finally, determine the meta-path hub type ψ P , divide P into a forward sub-path and a reverse sub-path The process is as Figure 4 shown. Among them, calculate the average connection density between target vertices Select D max corresponding meta-path.
[0067] S2. Construct a meta-path conversion graph by using the bidirectional meta-path instance search method according to the target meta-path.
[0068] Specifically, it includes the following steps:
[0069] S201. Determine the node type with the least number of vertices in the target meta-path as the split point type;
[0070] S202. Perform depth-first search along the forward and reverse directions of the target meta-path respectively, and collect the reachable target type node sets. Among them, define the forward direction as the sequential direction from the split point type of the target meta-path to the ending vertex type, and the reverse direction is the reverse order direction from the split point type to the starting vertex type;
[0071] S203. Perform a Cartesian product connection on the forward target type node set and the reverse target type node set to generate a complete meta-path conversion graph.
[0072] As described in the above steps S201 - S203, based on Figure 4 , for each vertex u of type ψ P , perform a depth - first search along the direction to obtain the reachable target vertex set Search along the direction to obtain the set Generate the meta - path instance edge set through the Cartesian product; construct the conversion graph adjacency matrix G in compressed sparse row format P , which supports neighbor queries with O(1) complexity; the constructed isomorphic graph is as Figure 3 shown.
[0073] S3. Implement non - independent and independent dual - mode clustering models on the meta - path conversion graph to obtain the core clustering results; among them, for the independent clustering model, use recursive - partition - based independent path verification to ensure the complete independence between path instances, and eliminate redundant calculations through the boundary pruning strategy.
[0074] Specifically, it includes the following steps:
[0075] S301. Implement the non - independent clustering model on the meta - path conversion graph to obtain the first core clustering result. Specifically, it includes the following steps:
[0076] S3011. Define the structural similarity between nodes in the non - independent mode. For any two nodes u and v in the meta - path conversion graph, their structural similarity σ P (u, v) is defined as the ratio of the number of their common neighbors to the geometric mean of the number of their neighbors, and the calculation formula is: where σ P (u, v) represents the similarity between nodes u and v, and N P (u) represents the set of neighbor nodes of node u based on the meta - path P.
[0077] S3012. Identify core nodes by setting the similarity threshold ∈ and the minimum number of neighbors μ. If the ∈ - neighbor number of a node u is greater than or equal to μ, then u is defined as a core node; among them, the ∈ - neighbor number of the node u is the number of neighbors whose structural similarity with u is greater than or equal to ∈.
[0078] S3013. Continuously expand the clustering through the core nodes until no further expansion is possible to obtain the first core clustering result.
[0079] S302. Implement the independent clustering model on the meta - path conversion graph to obtain the second core clustering result.
[0080] The independent clustering model process adopts the overall framework of the non - independent model and makes the following improvements:
[0081] (1) The independent - mode clustering model introduces an independent path verification mechanism based on recursive segmentation to ensure that the connection paths between nodes do not share intermediate entities; the independent path verification mechanism first recursively divides the target meta - path into a left - hand path P l and a right - hand path P r . Then, for each candidate common node w, it calculates the set of left - hand path instances from u to w and the set of right - hand path instances from v to w respectively. Next, it constructs a candidate node matching matrix, and enumerates the path combinations that do not share intermediate nodes through backtracking. Finally, when there are three completely independent groups of path instances of (u, w), (v, w), and (u, v), it determines that w is an independent common neighbor.
[0082] The specific process of the independent path verification algorithm is as follows: input the candidate node triple (u, v, w); recursively divide the meta - path until the sub - path length ≤ 2; generate a set of candidate intermediate nodes N cand for each segmentation level; traverse N cand through backtracking to ensure that the nodes at each level are not repeated; return verification success when there is a complete independent path chain.
[0083] (2) The independent - mode clustering model introduces an independence constraint in the calculation of structural similarity: where represents the independent path similarity between nodes u and v, and I P (u, v) represents the set of common neighbors connected by independent meta - path instances between u and v.
[0084] (3) During the clustering process of the independent - mode clustering model, it implements a boundary pruning strategy through boundary maintenance to eliminate redundant calculations. The process is as follows: pre - calculate the core candidate set in the non - independent mode as the upper bound based on pSCAN; maintain the upper and lower bounds (lb, ub) of the number of similar neighbors of each node; where: lb(u) is the number of confirmed neighbors, and ub(u) is the number of neighbors with σ P (u, v) ≥ ε in the non - independent mode; directly prune non - core nodes when ub < μ; implement a delayed verification strategy for boundary nodes, and only perform a complete independence check when the number of un - verified neighbors of lb + ≥ μ.
[0085] S4. Identify hub nodes and outliers based on the obtained core clustering results. The hub nodes need to meet the cross - cluster adjacency condition, and the outliers need to be excluded from the core extension path. Construct a multi - dimensional node role feature vector to achieve visual annotation and complete the structural clustering of the heterogeneous information network.
[0086] Specifically, it includes the following steps:
[0087] S401. Identify hub nodes and outliers based on the obtained first core clustering result and second core clustering result;
[0088] Identifying hub nodes: A hub node is a node that connects multiple core clusters but does not belong to any core cluster. Its role in the graph is similar to a "bridge", and the process of connecting different core clusters is as follows: After completing the construction of the core clusters, traverse all the nodes that have not been assigned to the core clustering results, and then check whether these nodes are connected to the nodes in multiple core clustering results. If so, mark this node as a hub node;
[0089] Identifying outliers: Outliers are those nodes that neither belong to any core cluster nor act as hub nodes. They are usually nodes that are weakly connected or isolated from other nodes in the graph. The process is as follows: After completing the construction of the core clusters, traverse all the nodes that have not been assigned to the core clustering results, and then check whether these nodes are connected to the nodes in multiple core clustering results. If a node neither meets the conditions of a hub node nor belongs to any core clustering result, mark it as an outlier;
[0090] S402. Construct a multi-dimensional node role feature vector to achieve visual annotation and complete the structure clustering of the heterogeneous information network.
[0091] The following conditions need to be met during annotation:
[0092] (1) The core node needs to satisfy |N ε [u]| ≥ μ and be reachable by at least two core nodes;
[0093] (2) The hub node needs to be adjacent to ≥ 2 different clusters and the proportion of cross-cluster edges > θ;
[0094] (3) The outlier needs to be excluded from the core expansion path and the number of single-cluster adjacent edges < δ.
[0095] The rules for setting the role determination thresholds include: the cross-cluster edge ratio θ = 0.3 × the average number of connections between clusters; the outlier determination threshold δ = 0.1 × the average number of connections within a cluster; the core expansion path length limit ≤ 3 hops.
[0096] The specific examples of the present invention are as follows:
[0097] S1. Heterogeneous Network Meta-Path Modeling: The data source is a large-scale heterogeneous network containing multiple entity types, such as the DBLP academic network (authors, papers, conferences, topics), the IMDB movie network (actors, movies, genres), etc. Each entity type contains attribute features, such as the research direction of an author, the publication year of a paper, etc. Edge types represent the relationships between entities, such as the writing relationship between an author and a paper, the publication relationship between a paper and a conference, etc.
[0098] Specific Implementation of Meta-Path Selection: (1) Traverse the network schema to generate candidate meta-paths, such as APCPA (author - paper - conference - paper - author), AMCMA (actor - movie - genre - movie - actor), etc.; (2) Count the number of instances of each meta-path and retain the top 20% of the high-frequency paths; (3) Calculate the average connection density between target vertices Select D max corresponding meta-path; (4) Determine the hub type ψ P as the type with the fewest vertices in the path.
[0099] S2. Construction of Bidirectional Search Transformation Graph. The specific process of constructing the transformation graph is as follows: (1) For each ψ P type vertex u, perform a depth-first search along the direction to obtain the set of reachable target vertices (2) Search along the direction to obtain the set Generate the meta-path instance edge set through the Cartesian product; (3) Store the adjacency matrix in the Compressed Sparse Row (CSR) format to support neighbor queries with O(1) complexity; (4) Construction example: For the DBLP network, select the APCPA meta-path, use the conference type as the hub, and construct the collaboration network between authors.
[0100] S3. Identification of Core Nodes in the Dual-Mode Meta-Path Transformation Graph, Structure Similarity Calculation:
[0101] (1) The formula for the non-independent mode is: where σ P (u, v) represents the similarity between nodes u and v, and N P (u) represents the set of neighbor nodes of node u based on the meta-path P.
[0102] (2) The formula for the independent mode is: where represents the independent path similarity between nodes u and v, and I P (u, v) is the set of independent common neighbors.
[0103] (3) The independent mode adopts dynamic boundary maintenance: pre-compute the non-independent core candidate set as the upper bound; maintain the upper and lower bounds (lb, ub) of the number of similar neighbors for each node; directly prune non-core nodes when ub(u) < μ; implement a delayed verification strategy for boundary nodes, and perform a complete independence check only when lb + the number of un-verified neighbors ≥ μ.
[0104] (4) The independent mode adopts independent path verification based on recursive partitioning. The verification algorithm process is as follows: 1> Input the candidate node triple (u, v, w); 2> Recursively partition the meta-path until the sub-path length ≤ 2; 3> Generate a candidate intermediate node set N for each partitioning level cand ; 4> Traverse N through backtracking cand , ensuring that nodes at each level are not repeated; 5> When there is a complete independent path chain, return verification success, and the process is as Figure 5 shown.
[0105] S4. Perform role annotation on the verified meta-path conversion graph, as Figure 6 gives a display of a clustering result. The role determination rules are as follows:
[0106] Core node: |N ε [u]| ≥ μ and is reachable by at least two core nodes.
[0107] Hub node: Adjacent to ≥ 2 different clusters and the proportion of cross-cluster edges > θ, where θ = 0.3 × the average number of connections between clusters.
[0108] Outlier: Excluded from the core expansion path and the number of single-cluster adjacent edges < δ, where δ = 0.1 × the average number of connections within a cluster.
[0109] The present invention combines meta-path semantic constraints with structural similarity calculation, and constructs two heterogeneous network structure clustering models, namely independent and non-independent models, which significantly improves the accuracy of heterogeneous network clustering. While maintaining the hub recognition ability of the SCAN algorithm, it supports minute-level processing of networks on the scale of billions, and can be widely applied to the fields of social network analysis, knowledge graph mining, and recommendation system optimization.
[0110] As Figure 7 shown, the present invention also provides an efficient heterogeneous information network graph structured clustering system based on meta-path enhancement, including:
[0111] An acquisition module 1, used to acquire and parse multi-type nodes in the heterogeneous network, define a vertex type set and an edge type set to generate a candidate meta-path set, and select a target meta-path based on the target clustering vertices;
[0112] A construction module 2, used to construct a meta-path conversion graph by adopting a two-way meta-path instance search method according to the target meta-path;
[0113] A dual-mode clustering module 3 is used to implement a non-independent and independent dual-mode clustering model on the meta-path conversion graph to obtain a core clustering result. Among them, for the independent clustering model, independent path verification based on recursive segmentation is adopted to ensure the complete independence between path instances, and redundant calculations are eliminated through a boundary pruning strategy;
[0114] An identification and annotation module 4 is applied to identify hub nodes and outliers based on the obtained core clustering result. The hub nodes need to meet the cross-cluster adjacency condition, and the outliers need to be excluded from the core extended path. A multi-dimensional node role feature vector is constructed to achieve visual annotation, and the structural clustering of the heterogeneous information network is completed.
[0115] In one embodiment, the acquisition module 1 specifically includes:
[0116] A loading unit is used to load a heterogeneous information network from a data source. The heterogeneous information network contains various types of nodes and edges. By parsing the network structure, each type of node and its attribute information are extracted;
[0117] A meta-path set generation unit is used to define a vertex type set and an edge type set according to the node types and edge types in the heterogeneous information network. The vertex type set contains all different types of nodes in the network, and the edge type set contains all different types of edges. Based on the vertex type set and the edge type set, all possible meta-path sets are generated by combining different types of vertices and edges. Among them, a meta-path is a sequence of vertex types connected in a specific order and is used to represent the semantic relationship between nodes;
[0118] A screening unit is used to determine the target clustering vertex type, that is, the node type that needs to perform clustering analysis. Meta-paths with the same starting vertex type and ending vertex type as the target clustering vertex type are screened out from the generated meta-path sets. A candidate meta-path set is generated through network pattern traversal, and high-frequency meta-paths are screened out. Finally, the meta-path that maximizes the average connection density between target type nodes is selected as the target meta-path for subsequent clustering analysis tasks.
[0119] In one embodiment, the construction module 2 specifically includes:
[0120] A determination unit is used to determine the node type with the least number of vertices in the target meta-path as the segmentation point type;
[0121] A search unit is used to perform depth-first search in both the forward and reverse directions along the target meta-path to collect the reachable target type node sets. Among them, the forward direction is defined as the sequential direction from the segmentation point type of the target meta-path to the ending vertex type, and the reverse direction is the reverse order direction from the segmentation point type to the starting vertex type;
[0122] A meta-path conversion graph generation unit, which is used to perform a Cartesian product connection on the forward target type node set and the reverse target type node set to generate a complete meta-path conversion graph.
[0123] In one embodiment, the dual-mode clustering module 3 specifically includes:
[0124] A first clustering unit, which is used to implement a non-independent clustering model on the meta-path conversion graph to obtain a first core clustering result;
[0125] A second clustering unit, which is used to implement an independent clustering model on the meta-path conversion graph to obtain a second core clustering result.
[0126] In one embodiment, the first clustering unit specifically includes:
[0127] A definition subunit, which is used to define the structural similarity between nodes in the non-independent mode. For any two nodes u and v in the meta-path conversion graph, their structural similarity σ P (u, v) is defined as the ratio of their common neighbor number to the geometric mean of their neighbor numbers, and the calculation formula is: where σ P (u, v) represents the similarity between nodes u and v, and N P (u) represents the set of neighbor nodes of node u based on the meta-path P;
[0128] An identification subunit, which is used to identify core nodes by setting a similarity threshold ∈ and a minimum neighbor number μ. If the ∈-neighbor number of a node u is greater than or equal to μ, then u is defined as a core node; where the ∈-neighbor number of the node u is the number of neighbors whose structural similarity to u is greater than or equal to ∈;
[0129] An extended clustering subunit, which is used to continuously perform extended clustering through core nodes until no further extension is possible to obtain a first core clustering result.
[0130] In one embodiment, in the second clustering unit, the independent clustering model process adopts the overall framework of the non-independent model and makes the following improvements:
[0131] The independent mode clustering model introduces an independent path verification mechanism based on recursive segmentation to ensure that the connection paths between nodes do not share intermediate entities; the independent path verification mechanism first recursively divides the target meta-path into a left half path P l and a right half path P r, then for each candidate common node w, calculate the set of left - half path instances from u to w and the set of right - half path instances from v to w respectively. Then construct a candidate node matching matrix, and enumerate path combinations without sharing intermediate nodes through backtracking. Finally, when there exist three groups of completely independent path instances of (u, w), (v, w), and (u, v), determine that w is an independent common neighbor;
[0132] The structural similarity calculation of the independent - mode clustering model introduces an independence constraint: where represents the independent path similarity between nodes u and v, and I P (u, v) represents the set of common neighbors connected by independent meta - path instances between u and v;
[0133] During the clustering process of the independent - mode clustering model, a boundary pruning strategy is implemented through boundary maintenance to eliminate redundant calculations. The process is as follows: pre - calculate the core candidate set in the non - independent mode as the upper bound; maintain the upper and lower bounds (lv, ub) of the number of similar neighbors of each node; directly prune non - core nodes when ub < μ; implement a delayed verification strategy for boundary nodes, and only perform a complete independence check when the number of un - verified neighbors of lb + ≥ u.
[0134] In one embodiment, the recognition and annotation module 4 specifically includes:
[0135] A recognition unit, configured to recognize hub nodes and outliers based on the obtained first core clustering result and second core clustering result;
[0136] Recognize hub nodes: After completing the construction of core clustering, traverse all nodes that have not been assigned to the core clustering result, and then check whether these nodes are connected to nodes in multiple core clustering results. If so, mark this node as a hub node;
[0137] Recognize outliers: After completing the construction of core clustering, traverse all nodes that have not been assigned to the core clustering result, and then check whether these nodes are connected to nodes in multiple core clustering results. If a node neither meets the conditions of a hub node nor belongs to any core clustering result, mark it as an outlier;
[0138] An annotation unit, configured to construct a multi - dimensional node role feature vector to achieve visual annotation and complete the structural clustering of the heterogeneous information network.
[0139] Each of the above - mentioned modules, units, and subunits is used to correspondingly execute each step in the above - mentioned efficient heterogeneous information network graph - structured clustering method based on meta - path enhancement. The specific implementation manner refers to the method embodiments described above and will not be elaborated here.
[0140] Such asFigure 8 As shown, the present invention also provides a computer device, which may be a server, and its internal structure may be as Figure 8 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store all the data required for the process of the efficient heterogeneous information network graph structured clustering method based on meta-path enhancement. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the efficient heterogeneous information network graph structured clustering method based on meta-path enhancement.
[0141] Those skilled in the art can understand that Figure 8 the structure shown in
[0142] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied.
[0143] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0144] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, apparatus, article, or method that includes the element.
[0145] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. An efficient heterogeneous information network graph structured clustering method based on meta-path enhancement, characterized in that, Including: S1. Obtain and parse multi-type nodes in a heterogeneous network, define a vertex type set and an edge type set to generate a candidate meta-path set, and select a target meta-path based on target clustering vertices; S2. Construct a meta-path conversion graph by adopting a two-way meta-path instance search method according to the target meta-path; S3. Implement a non-independent and independent dual-mode clustering model on the meta-path conversion graph to obtain a core clustering result; among them, for the independent clustering model, independent path verification based on recursive segmentation is adopted to ensure the complete independence between path instances, and redundant calculations are eliminated through a boundary pruning strategy; S4. Identify hub nodes and outliers based on the obtained core clustering result. The hub nodes need to meet the cross-cluster adjacency condition, and the outliers need to be excluded from the core extended path. Construct a multi-dimensional node role feature vector for visual annotation to complete the structure clustering of the heterogeneous information network.
2. The efficient heterogeneous information network graph structured clustering method based on meta-path enhancement according to claim 1, wherein The specific steps of step S1 include: S101. Load a heterogeneous information network from a data source. The heterogeneous information network contains multiple types of nodes and edges. By parsing the network structure, extract each type of node and its attribute information; S102. According to the node types and edge types in the heterogeneous information network, define a vertex type set and an edge type set. The vertex type set contains all different types of nodes in the network, and the edge type set contains all different types of edges. Based on the vertex type set and the edge type set, generate all possible meta-path sets by combining different types of vertices and edges; among them, a meta-path is a sequence of vertex types connected in a specific order, which is used to represent the semantic relationship between nodes; S103. Determine the target clustering vertex type, that is, the node type that needs to be clustered. Screen out the meta-paths whose starting vertex type and ending vertex type are the same as the target clustering vertex type from the generated meta-path set. Generate a candidate meta-path set through network pattern traversal, and screen out high-frequency meta-paths. Finally, select the meta-path with the largest average connection density between target type nodes as the target meta-path for subsequent clustering analysis tasks.
3. The efficient heterogeneous information network graph structured clustering method based on meta-path enhancement according to claim 1, characterized in that The specific steps of step S2 include: S201. Determine the node type with the least number of vertices in the target meta-path as the split point type; S202. Perform depth-first search along the forward and reverse directions of the target meta-path respectively, and collect the reachable target type node sets; among them, the forward direction is defined as the sequential direction from the split point type of the target meta-path to the ending vertex type, and the reverse direction is the reverse order direction from the split point type to the starting vertex type; S203. Perform a Cartesian product connection on the forward target type node set and the reverse target type node set to generate a complete meta-path conversion graph.
4. The efficient heterogeneous information network graph structured clustering method based on meta-path enhancement according to claim 1, wherein The specific steps of step S3 include: S301. Implement a non-independent clustering model on the meta-path conversion graph to obtain a first core clustering result; S302. Implement an independent clustering model on the meta-path conversion graph to obtain a second core clustering result.
5. The efficient heterogeneous information network graph structured clustering method based on meta-path enhancement according to claim 4, characterized in that The specific steps of step S301 include: S3011. Define the structural similarity between nodes in the non - independent mode. For any two nodes u and v in the meta - path conversion graph, their structural similarity σ P (u, v) is defined as the ratio of the number of their common neighbors to the geometric mean of the number of their neighbors, and the calculation formula is: where σ P (u, v) represents the similarity between nodes u and v, and N P (u) represents the set of neighbor nodes of node u based on the meta - path P; S3012. Identify core nodes by setting a similarity threshold ∈ and a minimum number of neighbors μ. If the ∈-neighbor number of a node u is greater than or equal to μ, then u is defined as a core node; where the ∈-neighbor number of the node u is the number of neighbors whose structural similarity to u is greater than or equal to ∈. S3013. Continuously expand and cluster through the core nodes until no further expansion is possible, obtaining the first core clustering result.
6. The efficient heterogeneous information network graph structured clustering method based on meta-path enhancement according to claim 5, characterized in that, In step S302, the independent clustering model process adopts the overall framework of the non-independent model and makes the following improvements: The independent mode clustering model introduces an independent path verification mechanism based on recursive segmentation to ensure that the connection paths between nodes do not share intermediate entities; the independent path verification mechanism first recursively segments the target meta-path into a left half-path P l and a right half-path P r . Then, for each candidate common node w, it calculates the set of left half-path instances from u to w and the set of right half-path instances from v to w respectively. Next, it constructs a candidate node matching matrix and enumerates path combinations that do not share intermediate nodes through backtracking. Finally, when there are three groups of completely independent path instances of (u, w), (v, w), and (u, v), it determines that w is an independent common neighbor; The structure similarity calculation of the independent mode clustering model introduces independence constraints: where represents the independent path similarity between nodes u and v, and I P (u, v) represents the set of common neighbors connected by independent meta-path instances between u and v; During the clustering process, the independent mode clustering model eliminates redundant calculations by implementing a boundary pruning strategy through boundary maintenance. The process is as follows: pre-compute the core candidate set in the non-independent mode as the upper bound; maintain the upper and lower bounds (lb, ub) of the number of similar neighbors for each node; directly prune non-core nodes when ub < μ; implement a delayed verification strategy for boundary nodes, and only perform a complete independence check when the number of unverified neighbors of lb + ≥ μ 7. The efficient heterogeneous information network graph structured clustering method based on meta-path enhancement according to claim 6, characterized in that, Step S4 specifically includes: S401. Identify hub nodes and outliers based on the obtained first core clustering result and the second core clustering result. Identify hub nodes: After completing the construction of the core clustering, traverse all the nodes that have not been assigned to the core clustering result, and then check whether these nodes are connected to the nodes in multiple core clustering results. If so, mark this node as a hub node. Identify outliers: After completing the construction of the core clustering, traverse all the nodes that have not been assigned to the core clustering result, and then check whether these nodes are connected to the nodes in multiple core clustering results. If a node neither meets the conditions of a hub node nor belongs to any core clustering result, mark it as an outlier. S402. Construct a multi-dimensional node role feature vector to achieve visual annotation and complete the structural clustering of the heterogeneous information network.
8. An efficient heterogeneous information network graph structured clustering system based on meta-path enhancement, characterized in that, Including: An acquisition module for acquiring and parsing multi-type nodes in the heterogeneous network, defining a vertex type set and an edge type set to generate a candidate meta-path set, and selecting a target meta-path based on the target clustering vertices. A construction module for constructing a meta-path conversion graph by means of bidirectional meta-path instance search according to the target meta-path. A dual-mode clustering module for implementing non-independent and independent dual-mode clustering models on the meta-path conversion graph to obtain a core clustering result; where for the independent clustering model, recursive segmentation-based independent path verification is adopted to ensure the complete independence between path instances, and redundant calculations are eliminated through a boundary pruning strategy. An identification and annotation module for identifying hub nodes and outliers based on the obtained core clustering result. The hub node needs to meet the cross-cluster adjacency condition, and the outlier needs to be excluded from the core expansion path. Construct a multi-dimensional node role feature vector to achieve visual annotation and complete the structural clustering of the heterogeneous information network.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.