A Heterogeneous Network Community Discovery Method and System
By calculating the centrality and similarity of the interaction chain in a heterogeneous network, selecting seed nodes and using seed diffusion algorithm to detect the community structure, the problem of difficulty in effectively detecting heterogeneous network communities in the existing technology is solved, and stable and effective community detection and overlapping community detection are achieved.
Patent Information
- Application Number
- CN202111499733.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-09
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-12-09
AI Technical Summary
Existing community detection algorithms are difficult to effectively deal with heterogeneity of heterogeneous networks, resulting in distortion of detection results or difficulty in obtaining unique detection results.
Starting from the interactive chain of the heterogeneous network, the centrality of the interaction chain and similarity of each node are calculated, and the node with the largest centrality of the interaction chain is selected as the seed node, and the community structure is detected using the seed diffusion algorithm and the label propagation algorithm.
Effectively deal with the heterogeneity of heterogeneous networks, detect the community structure that conforms to the practical significance of heterogeneous networks, avoid the randomness of traditional algorithms, and can detect overlapping communities.
Smart Images

Figure CN114283021B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of network technology, and more specifically, relates to a method and system for discovering heterogeneous network communities. Background Art
[0002] Since the discovery of scale-free networks and small-world networks, complex networks have been a hot topic in academic research. Abstracting complex systems as complex networks and studying them can help people deeply understand the characteristics of complex systems and guide people to optimize, control, and use complex systems. In traditional research, complex systems are usually abstracted as homogeneous complex networks, that is, both nodes and edges have the same attributes. The research on homogeneous complex networks is currently relatively extensive, and its processing method is also relatively convenient. However, with the in-depth research, it is found that abstracting complex systems as homogeneous networks does not conform to the facts in many cases, and this abstraction is too simple. For example, in a citation network, nodes are divided into three categories: authors, articles, and journals. Therefore, many scholars have begun to study heterogeneous networks to discover the properties of real complex systems.
[0003] Community structure is an important structure evolved from complex networks. Community structure usually has the following characteristics: the connections between nodes within a community are relatively close, and the connections between communities are relatively loose. Discovering and studying community structure can well solve many problems in reality. For example, studying the community structure in a social network can discover groups of people with the same interests; studying the community structure in a protein network can discover proteins with the same functions. Therefore, many scholars have proposed various community detection algorithms to ensure that the community structure in the network can be accurately and effectively detected. Most of these community detection algorithms are for homogeneous complex networks, that is, when detecting communities, the nodes and edges in the network are considered to be the same. This assumption can greatly reduce the difficulty of community detection algorithms and also help improve the efficiency of the algorithms.
[0004] Currently, there is little research on community detection in heterogeneous networks. Usually, the following two methods are adopted: one is to ignore the heterogeneity of network nodes and edges and directly use the community detection algorithm for homogeneous networks to detect communities; the other is to use a certain type of node as the seed node, and use the homogeneous community detection algorithm to detect the community of the seed node to generate a seed community, and then absorb other types of nodes into the seed community to obtain the final community structure of the entire network. There are problems with both of the above algorithms: the first method directly ignores the heterogeneity of the network, causing important information in the network to be lost and resulting in the distortion and invalidation of the detection results; the second method must select appropriate seed nodes. Different selections of seed nodes will lead to different final results of community detection, and it is difficult to obtain a unique detection result. Moreover, the detection method of first using seed nodes and then other nodes severs the connection between different types of nodes, and this connection is precisely the cause of the heterogeneity of some networks. In summary, the existing technical means are still difficult to effectively handle the community detection of complex heterogeneous networks. Summary of the Invention
[0005] In view of at least one defect or improvement requirement of the prior art, the present invention provides a method and system for discovering communities in heterogeneous networks. Starting from the interaction chains formed by heterogeneous networks, it can effectively handle the heterogeneity of heterogeneous networks and detect community structures that conform to the actual meaning of heterogeneous networks.
[0006] To achieve the above object, according to the first aspect of the present invention, a method for discovering communities in heterogeneous networks is provided, including the steps of:
[0007] Search for and record all interaction chains in the heterogeneous network. An interaction chain is a link formed by the interaction of each node in the network;
[0008] Calculate the interaction chain centrality of each node in the network. The interaction chain centrality is a value that describes the quantity and quality of the interaction chains passing through a certain node. Select the node with the largest interaction chain centrality in the area as the seed node. The largest interaction chain centrality in the area means that a certain node has the largest interaction chain centrality compared with the nodes connected to it;
[0009] Determine the label of the seed node, and let the seed node spread its own label to the nodes connected to it and meeting the preset conditions. The connected nodes that obtain the label then continue to expand their own labels until all nodes in the network obtain labels, and determine the community according to the labels of all nodes.
[0010] Further, the preset condition is the condition that the interaction chain centrality and the interaction chain similarity need to meet. The interaction chain similarity is a value that describes the situation of two nodes sharing interaction chains.
[0011] Further, the preset condition is that the interaction chain centrality of the seed node is greater than that of its connected nodes, and the interaction chain similarity between the seed node and its connected nodes is greater than a preset threshold.
[0012] Further, the calculation formula for the interaction chain centrality is:
[0013]
[0014] where c x is the interaction chain centrality of node x, is the j-th interaction chain passing through node x, is a function describing the quality of the interaction chain quality.
[0015] Further, the calculation formula for the interaction chain similarity is:
[0016]
[0017] where sim(x, y) is the interaction chain similarity between nodes x and y, is a function describing the quality of the interaction chain quality, is a function describing the quality of the interaction chain quality, is the j-th interaction chain passing through node x, is the j-th interaction chain passing through node y.
[0018] Further, the determining of the community according to the labels of all nodes is to classify the nodes with the same label into the same community.
[0019] According to the second aspect of the present invention, there is provided a heterogeneous network community discovery system, including:
[0020] An interaction chain determination module, configured to search for and record all interaction chains in the heterogeneous network, where the interaction chain is a link formed by the interaction of each node in the network;
[0021] A seed node determination module, configured to calculate the interaction chain centrality of each node in the network, where the interaction chain centrality is a value describing the number and quality of the interaction chains passing through a certain node, and select the node with the largest interaction chain centrality in the area as the seed node, and the largest interaction chain centrality in the area means that the interaction chain centrality of a certain node is the largest compared with its connected nodes;
[0022] A label determination module, configured to determine the label of the seed node, and spread its own label from the seed node to the nodes connected to it and meeting the preset conditions, and the connected nodes that obtain the label then continue to expand their own labels until all nodes in the network obtain labels, and determine the community according to the labels of all nodes.
[0023] Generally speaking, compared with the prior art, the present invention has the following beneficial effects:
[0024] (1) Starting from the interaction chains formed in the heterogeneous network, the present invention can effectively handle the heterogeneity of the heterogeneous network and detect the community structure that conforms to the actual meaning of the heterogeneous network.
[0025] (2) By combining the seed diffusion algorithm and the label propagation algorithm, the present invention can stably and effectively detect communities, avoiding the randomness of traditional label algorithms.
[0026] (3) When the seeds are diffused in the present invention, each node can receive multiple labels, enabling the detection of overlapping communities. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a schematic diagram of the interaction chain of an embodiment of the present invention;
[0028] Figure 2 is a flowchart of a method for discovering communities in a heterogeneous network according to an embodiment of the present invention;
[0029] Figure 3 is a comparison diagram of the effects of a method for discovering communities in a heterogeneous network according to an embodiment of the present invention and other existing algorithms. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0031] The embodiment of the present invention proposes the concept of a heterogeneous network interaction chain. The interaction chain is the link formed by the interaction of each node in the heterogeneous network and is the fundamental reason for the formation of network heterogeneity. The interaction chain is an important tool for studying the properties of heterogeneous networks.
[0032] The heterogeneity of a heterogeneous network includes the heterogeneity of nodes and the heterogeneity of edges. Among them, the heterogeneity of nodes is determined by the type of nodes, and the heterogeneity of edges is generated by the interaction of different nodes. In a heterogeneous network, this interaction is usually the reason for the generation of network heterogeneity. It is also because of the need for interaction that heterogeneous networks are generated. For example, in a citation network, author nodes, paper nodes, and journal nodes also form a tight link. Authors obtain the latest cutting-edge results from journal nodes, write papers, and then submit them to journals, forming a link starting from journal nodes, reaching author nodes, going to paper nodes, and finally reaching journal nodes again. The embodiment of the present invention defines these links as heterogeneous network interaction chains. Figure 1Shows the interaction chains in the citation network, where J represents journal nodes, A represents author nodes, and P represents paper nodes. It can be seen that the existence of the heterogeneous network is precisely to serve the interaction links, and the interaction links also meet the needs of node interactions, which is the fundamental source of the formation of network heterogeneity.
[0033] Furthermore, the embodiments of the present invention propose two important parameters: interaction chain centrality and interaction chain similarity.
[0034] Among them, the interaction chain centrality is a value that describes the quantity and quality of the interaction chains passing through a certain node in a heterogeneous network, and can be expressed as:
[0035]
[0036] Among them, c x is the interaction chain centrality of node x, is the j-th interaction chain passing through node x, is a function that describes the quality of the interaction chain. For different heterogeneous networks, its specific manifestation forms are different. For example, in the citation network, the quality of the journals and authors that make up the interaction chain usually reflects the quality of the interaction chain.
[0037] The interaction chain similarity is a value that describes the situation of two nodes sharing interaction chains. The higher the interaction chain similarity between two nodes, the more high-quality interaction chains the two nodes share, and the closer the relationship between the two nodes in the heterogeneous network. The interaction chain centrality can be expressed as:
[0038]
[0039] Among them, sim(x,y) is the interaction chain similarity between nodes x and y, is a function that describes the interaction chain quality, is a function that describes the interaction chain quality, is the j-th interaction chain passing through node x, is the j-th interaction chain passing through node y. The definition method of formula (2) emphasizes more on the dependency relationship of node y on node x. The larger sim(x,y) is, the more the interaction chains passed by node y are the same as those passed by node x, the greater the influence of node x on node y, and the more likely node y should be assigned to the same community as node x.
[0040] As Figure 2 shown, a method for discovering communities in a heterogeneous network according to an embodiment of the present invention includes the steps of:
[0041] S1, search for and record all the interaction chains in the heterogeneous network. The interaction chain is the link formed by the interaction of each node in the network.
[0042] The specific method for searching and recording all interaction chains in the network is as follows: The depth-first algorithm is used to search and record all interaction chains in the network. The process of the depth-first algorithm is to delve into every possible branch path of the heterogeneous network until no further progress can be made, and each node can only be visited once. Starting from the starting point of each possible interaction chain and using the depth-first algorithm, the interaction chain starting from this point can be obtained. After traversing all starting points, the interaction chains in the entire network can be obtained.
[0043] S2. Calculate the interaction chain centrality of each node in the network. The interaction chain centrality is a value that describes the quantity and quality of the interaction chains passing through a certain node. Select the node with the maximum interaction chain centrality within the area as the seed node. The maximum interaction chain centrality within the area means that the interaction chain centrality of a certain node is the largest compared to the nodes connected to it.
[0044] Specifically, count and record the interaction chains passed by each node in the network, and then calculate the interaction chain centrality c of each node according to formula (1). x For a certain node, if the interaction chain centrality of this node is the largest compared to each node connected to it, then this node is the seed node. Judge each node in the network to find all the seed nodes in the network.
[0045] S3. Determine the label of the seed node, and let the seed node spread its own label to the nodes connected to it and meeting the preset conditions. The connected nodes that obtain the label will continue to expand their own labels until all nodes in the network obtain the label. Determine the community according to the labels of all nodes.
[0046] The label of the seed node is the serial number of the node. All nodes in the network have a unique serial number, and the label of the seed node is the serial number of the seed node.
[0047] Furthermore, the preset conditions are the conditions that the interaction chain centrality and the interaction chain similarity need to meet.
[0048] Furthermore, the preset conditions are as follows:
[0049] c seed >c neighbor (3)
[0050] sim(seed, neighbor)>thershold (4)
[0051] where c seed is the interaction chain centrality of the seed node seed, and c seedLet \(C_{seed, neighbor}\) be the interaction chain centrality of the nodes \(neighbor\) connected to the seed node, and \(sim(seed, neighbor)\) be the similarity of the interaction chain between the seed node \(seed\) and the connected node \(neighbor\), and \(thershold\) be the preset threshold.
[0052] The above preset conditions indicate that the interaction chain centrality of the seed node is greater than that of the surrounding nodes, and the similarity of the interaction chain between the seed node and the neighbor node is greater than a certain threshold. Among them, equation (3) shows that the seed node has a strong control ability and influence on the surrounding nodes, and the surrounding nodes can be well controlled or influenced by the seed node. Equation (4) shows that the surrounding nodes share more interaction chains with the seed node, and the two nodes cooperate very closely in the network. When the seed diffuses, each node retains all the received labels.
[0053] The surrounding nodes that obtain the labels continue to diffuse the labels to the nodes around them according to equations (3) and (4) until all nodes in the network obtain the labels.
[0054] Then, the communities are determined according to the labels of all nodes. Nodes with the same label can be classified into the same community.
[0055] Furthermore, when the seed diffuses in this method, since there is no limit on the number of labels of the nodes, each node can receive multiple labels. Since nodes with the same label are classified into the same community, when a node has multiple labels, then the node belongs to multiple communities. Nodes with multiple labels are the overlapping parts of the communities. Therefore, some of the obtained communities are overlapping, that is, this algorithm can detect overlapping communities.
[0056] Apply the heterogeneous network community discovery method (referred to as the ILLPA algorithm) of the embodiment of the present invention and three existing algorithms LZLPA, SLPA, and Modularity to a certain heterogeneous network, and perform community detection on the network to verify the algorithm effect by comparison. As Figure 3 shown, Figure 3 (a) is a heterogeneous network. Due to the heterogeneity of the network, the S node in this figure, as the end of the information flow, cannot participate in the OODA loop. Therefore, from the actual operation of the network, it cannot cooperate with other nodes and should not be assigned to any community. Figure 3 (b) is the community discovery result of the heterogeneous network community discovery method of the embodiment of the present invention. It can be seen that this method can accurately identify that this node should belong to an independent node and does not belong to any community. Figure 3(c), (d), and (e) are the community discovery results of the other three algorithms, namely LZLPA, SLPA, and Modularity. It can be seen that none of them can discover that this node is an independent node and classifies this node into other communities. Therefore, the heterogeneous network community discovery method of the embodiment of the present invention can well handle the heterogeneity of the network and find the heterogeneous network community structure that conforms to the actual operation law of the network.
[0057] A heterogeneous network community discovery system according to an embodiment of the present invention includes:
[0058] An interaction chain determination module, configured to record all interaction chains in the heterogeneous network, where an interaction chain is a link formed by the interaction of each node in the network;
[0059] A seed node determination module, configured to calculate the interaction chain centrality of each node in the network, where the interaction chain centrality is a value describing the quantity and quality of the interaction chains passing through a certain node, and select the node with the largest interaction chain centrality in the area as the seed node. The largest interaction chain centrality in the area means that a certain node has the largest interaction chain centrality compared with the nodes connected to it;
[0060] A label determination module, configured to determine the label of the seed node, and the seed node spreads its own label to the nodes connected to it and meeting the preset conditions. The nodes connected to it that obtain the label then continue to expand their own labels until all nodes in the network obtain the label, and determine the community according to the labels of all nodes.
[0061] Further, the preset condition is the condition that the interaction chain centrality and the interaction chain similarity need to meet, and the interaction chain similarity is a value describing the situation of the interaction chains shared by two nodes.
[0062] Further, the preset condition is that the interaction chain centrality of the seed node is greater than the interaction chain centrality of its connected nodes, and the interaction chain similarity between the seed node and its connected nodes is greater than the preset threshold.
[0063] Further, the determining the community according to the labels of all nodes is to classify the nodes with the same label into the same community.
[0064] The implementation principle and technical effect of the system are similar to those of the above method and will not be elaborated here.
[0065] It must be noted that in any of the above embodiments, the method does not necessarily execute in the order of the serial numbers. As long as it cannot be inferred from the execution logic that it must be executed in a certain order, it means that it can be executed in any other possible order.
[0066] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A heterogeneous network community discovery method, characterized in that Including the steps: Search for and record all interaction chains in the heterogeneous network. An interaction chain is a link formed by the interaction of each node in the network. The interaction chain includes a link in the citation network that starts from a journal node, goes to an author node, then to a paper node, and finally reaches a journal node again; Calculate the interaction chain centrality of each node in the network. The interaction chain centrality is a value that describes the quantity and quality of the interaction chains passing through a certain node. Select the node with the largest interaction chain centrality in the area as the seed node. The largest interaction chain centrality in the area means that a certain node has the largest interaction chain centrality compared to the nodes connected to it. The quality of the interaction chain includes the quality of the journals and authors that make up the interaction chain in the citation network; Determine the label of the seed node, and let the seed node spread its own label to the nodes connected to it and meeting the preset conditions. The connected nodes that obtain the label then continue to expand their own labels until all nodes in the network obtain labels, and determine the communities based on the labels of all nodes; The preset condition is that the interaction chain centrality of the seed node is greater than the interaction chain centrality of its connected nodes, and the interaction chain similarity between the seed node and its connected nodes is greater than the preset threshold; The calculation formula for the interaction chain centrality is: The calculation formula for the interaction chain similarity is: Among them, is x the interaction chain centrality of the node, and x is y the interaction chain similarity of the node, which is obtained through x the j th interaction chain of the node, and is obtained through y the i th interaction chain of the node, is a function describing the quality of the interaction chain and is a function describing the quality of the interaction chain .
2. The heterogeneous network community discovery method according to claim 1, wherein The determination of the community based on the labels of all nodes is to group the nodes with the same label into the same community.
3. A heterogeneous network community discovery system, characterized in that, Including: An interaction chain determination module for searching for and recording all interaction chains in the heterogeneous network. An interaction chain is a link formed by the interaction of each node in the network. The interaction chain includes a link in the citation network that starts from a journal node, goes to an author node, then to a paper node, and finally reaches a journal node again; A seed node determination module for calculating the interaction chain centrality of each node in the network. The interaction chain centrality is a value that describes the quantity and quality of the interaction chains passing through a certain node. Select the node with the largest interaction chain centrality in the area as the seed node. The largest interaction chain centrality in the area means that a certain node has the largest interaction chain centrality compared to the nodes connected to it. The quality of the interaction chain includes the quality of the journals and authors that make up the interaction chain in the citation network; A label determination module for determining the label of the seed node, and letting the seed node spread its own label to the nodes connected to it and meeting the preset conditions. The connected nodes that obtain the label then continue to expand their own labels until all nodes in the network obtain labels, and determine the communities based on the labels of all nodes; The preset condition is that the interaction chain centrality of the seed node is greater than the interaction chain centrality of its connected nodes, and the interaction chain similarity between the seed node and its connected nodes is greater than the preset threshold; The calculation formula for the interaction chain centrality is: The calculation formula for the interaction chain similarity is: Among them, is x the interaction chain centrality of the node, and x is y the interaction chain similarity of the node, is calculated through x the j th interaction chain of the node, and y is calculated through i the th interaction chain of the node, is a function describing the quality of the interaction chain and is a function describing the quality of the interaction chain 4. The heterogeneous network community discovery system according to claim 3, wherein The determination of the community based on the labels of all nodes is to group the nodes with the same label into the same community.