Distributed community discovery method and system
Through the combination of distributed computing and attribute modularity, fast and accurate community division is achieved, the problem of slow community division in traditional methods is solved, and the efficiency and timeliness of risk warning are improved.
Patent Information
- Application Number
- CN202510121585.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-09
AI Technical Summary
Traditional community discovery methods are difficult to quickly divide communities, which affects the efficiency and timeliness of risk warning.
The distributed community discovery method is adopted to calculate the attribute modules of each community in the graph through distributed computing to achieve rapid community division. The specific steps include the master node sending community division proposals to the work nodes, the work node calculates the community embedding characterization and attribute module degree, and sends the results to the master node. The master node determines the global attribute module degree based on the received attribute module degree to evaluate the division quality.
It improves the accuracy and rationality of community division, and enhances the efficiency of large-scale relationship graph module calculations and the timeliness of result outputs.
Smart Images

Figure CN119961694A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of knowledge graphs, and in particular, to a distributed community discovery method and system. Background Art
[0002] Risk control is required in various business systems. For example, in e-commerce platforms, it is necessary to guard against the risk of gray market fraud, and in payment platforms, it is necessary to guard against risks such as fraudulent transactions and illegal fund transfers. Traditional risk assessment models mainly rely on individual operation data and often ignore important information contained in interpersonal networks. Research has also found that various risks in business platforms reflect the characteristics of joint operations by gangs. Therefore, identifying communities or groups composed of multiple users or user devices based on relationship graphs is of great significance for risk control.
[0003] However, as the scale of the relationship graph continues to expand, traditional community discovery methods are difficult to quickly divide communities, which may affect the efficiency and timeliness of risk warning. Therefore, a method is needed to more quickly discover communities based on relationship graphs. Summary of the invention
[0004] One or more embodiments of this specification describe a distributed community discovery method and system, which is based on the paradigm of distributed computing and calculates the attribute modularity of each community in a graph to perform community division more quickly.
[0005] In a first aspect, a distributed community discovery method is provided, which is applied to a distributed system, wherein the distributed system includes a master node and a plurality of worker nodes, each of which holds a graph shard of a full graph; the method includes:
[0006] The master node sends a first community division proposal for the entire graph to each worker node, which includes several communities.
[0007] Any target working node determines the target community embedding representation based on the embedding representation of each graph node attribute in any target community on the graph shard it holds, and then determines the target attribute modularity of the target community, and sends it to the master node; the target attribute modularity is negatively correlated to the first similarity between the target community embedding representation and the global embedding representation; the global embedding representation is determined based on the aggregation result of the embedding representation of each graph node attribute in the whole graph;
[0008] The master node determines the global attribute modularity according to the received attribute modularity of each community; the global attribute modularity is used to evaluate the partition quality of the first community partition proposal.
[0009] In some possible implementations, the first community division proposal is an update proposal determined for an existing community division result, wherein there is at least one migration node, the community to which the migration node belongs in the first community division proposal is different from the community to which the migration node belongs in the existing community division result.
[0010] In some possible implementations, the following further includes:
[0011] The master node moves at least one graph node in any community in the existing community division results to an adjacent community to obtain the first community division proposal.
[0012] In some possible implementations, the target community includes a target central graph node, and the target central graph node is located on a graph shard held by a target working node; and determining the target community embedding representation includes:
[0013] The target worker node receives the embedded representation of each graph node attribute from each worker node holding the graph node of the target community;
[0014] The target working node performs a first aggregation process on the embedded representations of the attributes of each graph node in the target community, thereby determining the embedded representation of the target community.
[0015] In some possible implementations, the global embedding representation is determined by the following steps:
[0016] Any working node determines the embedded representation of the shard based on the embedded representation of each graph node attribute on the graph shard it holds and sends it to the master node;
[0017] The master node performs a second aggregation process based on the received embedding representations of each shard, thereby determining a global embedding representation and sending it to each working node.
[0018] In some possible implementations, the target attribute modularity is positively correlated to the first target norm of the target community embedding representation.
[0019] In some possible implementations, determining the target attribute modularity of the target community includes:
[0020] Determine a first ratio according to a ratio of the square of the first similarity to a second target norm of the global embedding representation;
[0021] The target attribute modularity is determined according to a weighted difference between the first target norm and the first ratio.
[0022] In some possible implementations, determining the global attribute modularity includes:
[0023] The global attribute modularity is determined based on the ratio of the sum of the attribute modularity of each community to the second objective norm of the global embedding representation.
[0024] In some possible implementations, the following further includes:
[0025] The target working node determines the target structural modularity of the target community according to the relevant edge weights and global weight of the target community, and sends it to the master node; the global weight is determined by aggregating the edge weights of each relationship edge in the full graph;
[0026] The master node determines the global structural modularity according to the received structural modularity of each community; the global structural modularity is used to evaluate the partition quality of the first community partition proposal together with the global attribute modularity.
[0027] In some possible implementations, determining a target structural modularity of a target community includes:
[0028] The target working node aggregates the edge weights of the internal relationship edges of the target community to obtain the internal weight of the target community, and aggregates the edge weights of the relationship edges connected to each graph node in the target community to obtain the total weight of the target community;
[0029] The target working node calculates the target structure modularity according to the internal weight and total weight of the target community and the global weight, so that the modularity is positively correlated with the internal weight and negatively correlated with the total weight.
[0030] In some possible implementations, the target community includes a target central graph node, and the target central graph node is located on a graph shard held by a target working node; before determining the target structural modularity of the target community, the method further includes:
[0031] The target working node receives the edge weights of the relationship edges connecting each graph node from each working node holding the graph node of the target community.
[0032] In some possible implementations, the global weight is determined by the following steps:
[0033] Any working node determines the total node weight of the target graph node based on the sum of the edge weights of each relationship edge connected to any target graph node on the graph shard it holds, and sends it to the master node;
[0034] The master node determines the global weight based on the sum of the total node weights of each node received and sends it to each working node.
[0035] In some possible implementations, calculating the modularity of the target structure includes:
[0036] Determining a second ratio according to a ratio between an internal weight of the target community and a global weight;
[0037] Determining a third ratio according to a ratio between the total weight of the target community and the global weight;
[0038] The target structural modularity is determined according to a weighted difference between the second ratio and the square of the third ratio.
[0039] In some possible implementations, the full graph is a user relationship graph, in which graph nodes represent users, and relationship edges represent relationships between users; and the community represents a group composed of several users.
[0040] In the second aspect, a distributed community discovery system is provided, including a master node and a plurality of worker nodes, each of which holds a graph shard of the entire graph, wherein:
[0041] The master node is used to send the first community division proposal for the entire graph to each working node, which includes several communities;
[0042] Any target working node is used to determine the target community embedding representation based on the embedding representation of each graph node attribute in any target community on the held graph shard, and then determine the target attribute modularity of the target community, and send it to the master node; the target attribute modularity is negatively correlated to the first similarity between the target community embedding representation and the global embedding representation; the global embedding representation is determined based on the aggregation result of the embedding representation of each graph node attribute in the whole graph;
[0043] The master node is further used to determine the global attribute modularity according to the received attribute modularity of each community; the global attribute modularity is used to evaluate the partition quality of the first community partition proposal.
[0044] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method of the first aspect.
[0045] According to a fourth aspect, a computing device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method of the first aspect is implemented.
[0046] The distributed community discovery method and system proposed in the embodiments of this specification are applied to a distributed system. The master node of the distributed system proposes a community partition proposal and sends it to each working node of the distributed system. Each working node calculates the attribute modularity of the community on the graph shard it holds based on the community in the community partition proposal, and sends it to the master node. Then, the master node determines the global attribute modularity of the community partition proposal based on the attribute modularity of each community. The global attribute modularity can be used to evaluate the partition quality of the community partition proposal, and then a partition proposal with a high partition quality ranking can be selected from multiple community partition proposals and the partition can be executed. Based on the attribute modularity, the embodiments of this specification can improve the accuracy and rationality of community partitioning. At the same time, the use of distributed computing can improve the computational efficiency of the modularity of large-scale relationship graphs and the timeliness of the output results. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only the multiple embodiments disclosed in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0048] Figure 1 A schematic diagram showing an implementation scenario of a distributed community discovery method according to an embodiment;
[0049] Figure 2 A flow chart showing a distributed community discovery method according to one embodiment;
[0050] Figure 3 A schematic block diagram of a distributed community discovery system according to one embodiment is shown. DETAILED DESCRIPTION
[0051] The solution provided in this specification is described below in conjunction with the accompanying drawings.
[0052] As mentioned earlier, community discovery plays an important role in the field of risk control. Through community discovery, business platforms can identify users' potential positions and relationships in social networks, thereby assessing risks more comprehensively. For example, for a user relationship graph containing multiple users, it contains multiple communities, and each community contains at least one user. In the payment scenario, when multiple users in the same community have abnormal transaction behaviors, this may indicate that there is a systemic risk in the entire community. By analyzing the capital flow pattern, transaction frequency, and social relationship strength between community members, the payment platform can promptly discover potential fraudulent transaction behaviors and fraud risk transmission paths. Therefore, by continuously monitoring and analyzing changes in community structure, business platforms can build a more robust risk early warning system.
[0053] The community discovery methods in related technologies mainly use modularity as a standard to measure the quality of community division. However, this method has many shortcomings. On the one hand, modularity is calculated based on the structural relationship (edge weight) in the relationship graph, and does not take into account the differences in attributes between the various graph nodes in the relationship graph. In actual applications of relationship graphs, each graph node can represent different users, transactions, places, etc. If the differences between the various graph nodes in the relationship graph are not distinguished, a lot of key information will be lost. On the other hand, for application scenarios such as risk control that require high timeliness, considering the complexity of modularity calculation and the increasingly large data scale of the relationship graph, the traditional stand-alone computing model is difficult to support the demand for rapid output of community division results.
[0054] Based on the above analysis, the embodiments of this specification propose a distributed community discovery method, which introduces the attributes of graph nodes into the calculation process of modularity while improving the calculation efficiency of evaluation indicators related to community division quality, so as to improve the accuracy and rationality of community division.
[0055] First of all, it should be noted that the term "node" appears in the distributed system and the full graph of the embodiments of this specification. In the distributed system, it refers to the computing unit involved in distributed computing, which corresponds to the English word "node"; and in the full graph, it refers to the element of the graph, which corresponds to the English word "vertex". In order to distinguish the two, the embodiments of this specification refer to the "node" in the distributed system as the master node and the working node, and the "node" in the full graph and the graph shard as the graph node. If the word "node" appears alone, it refers to the graph node vertex.
[0056] Figure 1A schematic diagram of an implementation scenario of a distributed community discovery method according to an embodiment is shown. The method is applied to a distributed system. The distributed system includes a master node and several worker nodes, such as worker node 1 to worker node n. Each worker node holds a graph slice of the full graph. The graph slices of each worker node can be combined together to form the full graph. Figure 1 In the example, the relationship graph on the left may include the current community division result, which may be held by the master node and includes multiple communities, such as community 1 to community m. Each community is represented by a dotted arc circle, and the graph nodes inside an arc circle belong to the community represented by the arc circle. A community may include a graph node as the central node of the community. When the modularity is subsequently calculated, the working node holding the central node of the community is responsible for calculating the modularity of the community.
[0057] It should be noted that the process of community discovery can include multiple rounds of "proposing community division proposals - evaluating the quality of proposals - updating community division results". Figure 1 The following is a detailed description of the steps for evaluating the quality of a proposal in one round of the process. Initially, each graph node can belong to its own community, and then the above process is repeated for multiple rounds until the indicator for evaluating the quality of the proposal no longer changes or reaches the predetermined round limit, and then the final community division result is output.
[0058] First, the master node proposes a new community division proposal based on the current community division results, which includes moving at least one graph node from its community to another community. Figure 1 In the figure, the master node moves the gray graph nodes in the figure from the community below to the community on the right to form a new community division proposal, and sends this new community division proposal to each working node.
[0059] Then, each working node determines the community it is responsible for based on the community center nodes contained in the graph shard it holds, and then receives information about each graph node in each community it is responsible for from other working nodes if necessary. Based on this graph node information, the attribute modularity of each community is calculated and then sent to the master node. Unlike traditional modularity (hereinafter referred to as structural modularity), attribute modularity is obtained based on the attribute embedding representation of each graph node, and its detailed calculation method will be explained in detail in the subsequent steps.
[0060] Next, the master node determines the global attribute modularity based on the attribute modularity of each community received. The global attribute modularity is used to evaluate the quality of the new community partition proposal and determine whether to update the community partition result based on the community partition proposal. The details of evaluating the quality of the partition proposal will also be further explained in the subsequent steps.
[0061] It should be noted that Figure 1 It is merely a specific example, which does not constitute a limitation on the scope of protection of the embodiments of this specification.
[0062] In the above example description, the global attribute modularity is determined based on the attributes of each graph node in the graph and is used as one of the bases for community division. By using a distributed system to calculate the global attribute modularity of the partition proposal, community discovery can be performed more quickly and accurately.
[0063] The specific implementation steps of the above-mentioned distributed community discovery method are described below in conjunction with specific embodiments.
[0064] Figure 2 A flowchart of a distributed community discovery method according to an embodiment is shown. The execution subject of the method can be any platform or server or device cluster with computing and processing capabilities. Figure 2 As shown, the method is applied to a distributed system, which includes a master node and several working nodes, each working node holding a graph shard of the whole graph; the method at least includes: step 204, the master node sends a first community division proposal for the whole graph to each working node, which includes several communities; step 206, any target working node determines the target community embedding representation based on the embedding representation of each graph node attribute in any target community on the graph shard it holds, and then determines the target attribute modularity of the target community, and sends it to the master node; the target attribute modularity is negatively correlated to the first similarity between the target community embedding representation and the global embedding representation; the global embedding representation is determined based on the aggregation result of the embedding representation of each graph node attribute in the whole graph; step 208, the master node determines the global attribute modularity based on the attribute modularity of each community received; the global attribute modularity is used to evaluate the division quality of the first community division proposal.
[0065] It should be noted that the distributed system in the embodiments of this specification may include one or more working nodes. Figure 2 It only shows the interaction process between any one of the working nodes and the master node, which does not constitute a limitation on the number of working nodes in the distributed system.
[0066] The full graph contains multiple graph nodes and the relationship edges between graph nodes. The summary results of the graph shards of the full graph held by each working node can constitute the full graph itself. The process of dividing the full graph into multiple graph shards can be completed by the master node, or by the master node and the working node, which is not limited here. The partitioning method can use vertex partitioning or edge partitioning, which is not limited here.
[0067] In one embodiment, the aforementioned full graph may be a user relationship graph, in which graph nodes represent users, and relationship edges represent relationships between users; and the community represents a group composed of several users.
[0068] In different embodiments, the relationship between users represented by the relationship edge may be different specific relationships. For example, in a specific embodiment, the relationship between users includes one or more of transaction behaviors, fund transfer behaviors, and social behaviors between users.
[0069] In addition, each graph node may also include node attributes. In different embodiments, the content of the node attributes may be different, which is not limited here. For example, in a specific embodiment, the node attributes may include one or more of the account information, login time, and transaction frequency of the corresponding user of the graph node.
[0070] The specific execution process of each of the above steps is described below.
[0071] First, in step 204, the master node sends a first community division proposal for the entire graph to each working node, which includes several communities.
[0072] The first community division proposal may be a community division result for the entire graph, including several communities, each community including at least one graph node, and the summary result of the graph nodes in each community may be all the graph nodes in the entire graph. At the same time, each graph node belongs to only one community.
[0073] The first community division proposal can be obtained by directly dividing the original full graph, or can be obtained by modifying an existing community division result. The existing community division result can be the current community division result at the beginning of any round of multiple updates in the community discovery process.
[0074] In one embodiment, the first community division proposal is an update proposal determined for an existing community division result, wherein there is at least one migration node, the community to which the migration node belongs in the first community division proposal is different from the community to which the migration node belongs in the existing community division result.
[0075] A migration node is a node that is moved by the master node of the distributed system in the existing community division result. It is moved from the first community to which it belongs in the community division result to a second community that is different from the first community. The community division result after moving the migration node can be the first community division proposal.
[0076] Based on this, in a more specific embodiment, before step 204, the method further includes step 202, where the master node moves at least one graph node in any community in the existing community division results to an adjacent community to obtain the first community division proposal.
[0077] The adjacent community of a community where any graph node is located may be the community to which the neighboring nodes of the graph node in the entire graph belong.
[0078] Then, in step 206, any target working node determines the target community embedding representation Z according to the embedding representation z of each graph node attribute in any target community c on the graph shard it holds. c , and then determine the target attribute modularity of the target community And sent to the master node; the target attribute module Negatively correlated with the target community embedding representation Z c The first similarity sim(Z c ,Z); the global embedding representation Z is determined by the aggregation result of the embedding representation z of each graph node attribute in the whole graph. Wherein, sim() represents the similarity function.
[0079] The embedded representation of graph node attributes can be directly converted from the attribute information of the graph nodes, or a pre-trained text encoder such as BERT (Bidirectional Encoder Representations from Transformers) can be used to encode the attributes of each graph node. It can also be obtained by using knowledge graph representation learning or graph neural network calculation, which is not limited here.
[0080] Any target community on the graph shard held by the target working node in step 206 indicates that the target community is calculated by the target working node. Since each working node needs to be responsible for calculating the community embedding representations of several communities respectively, it is necessary to determine which working node is responsible for calculating each community in the first community partition proposal. In one embodiment, a central graph node can be set for each community. For any community, the working node that holds the central graph node in the graph shard is responsible for calculating the community. In another embodiment, the number of graph nodes held by each working node in each community can also be counted separately for each community. For any community, the working node that holds the largest number of graph nodes in the community is responsible for calculating the community.
[0081] In one embodiment, the target community includes a target central graph node, and the target central graph node is located on a graph slice held by a target working node. In this embodiment, step 206 determines the target community embedding representation Z c ,include:
[0082] The target working node receives the embedded representation of each graph node attribute from each working node holding the graph node of the target community; then, the target working node performs a first aggregation process on the embedded representation of each graph node attribute in the target community, thereby determining the target community embedded representation Z c .
[0083] The first aggregation process can be implemented in a variety of specific operations, which are not limited here. In one embodiment, the target working node sums the embedded representations of the attributes of each graph node in the target community to obtain the target community embedded representation Z c In another embodiment, the target working node averages the embedding representations of the attributes of each graph node in the target community to obtain the target community embedding representation Z c .
[0084] The target working node collects the attribute embedding representation of each graph node in the target community from each working node, and then performs the first aggregation process on it to obtain the target community embedding representation Z of the target community. c .
[0085] In one embodiment, the global embedding representation Z is determined by the following steps:
[0086] Any working node determines the shard embedding representation based on the embedding representation of each graph node attribute on the graph shard it holds and sends it to the master node; the master node performs a second aggregation process based on the received shard embedding representations to determine the global embedding representation Z and sends it to each working node.
[0087] The shard embedding representation can be determined based on the sum or average of the embedding representations of the attributes of each graph node on the graph shard, which is not limited here.
[0088] The second aggregation process can be implemented in a variety of specific operations, which are not limited here. In one embodiment, the master node sums the received embedding representations of each shard to obtain a global embedding representation Z. In another embodiment, the master node averages the received embedding representations of each shard to obtain a global embedding representation Z.
[0089] Since the global embedding representation Z is not affected by the community division result, the value of the global embedding representation Z is always unchanged when the entire graph remains unchanged, so the above steps can be pre-executed before step 204. After receiving the global embedding representation Z, each working node can save it and use it in step 206 without recalculating it in each round of community division.
[0090] After getting the target community embedding representation Z c Afterwards, the target attribute modularity of the target community can be determined based on it and the pre-determined and saved global embedding representation Z The target attribute modularity Negatively correlated with the target community embedding representation Z c The first similarity sim(Z c ,Z c ). That is, the target community embedding representation Z c The smaller the similarity with the global embedding representation Z, the higher the modularity of the target attribute. The larger it is, the higher the division quality of the first community division proposal is.
[0091] In one embodiment, the target attribute modularity Positively correlated with the target community embedding representation Z c The first objective norm ‖Z c ‖. The value of the first target norm can be determined according to any norm, which is used to measure the size of a vector, such as 1-norm, 2-norm, infinity norm, etc., which are not limited here. That is, the larger the "size" of the target community embedding representation itself, the higher the modularity of the target attribute. The larger it is, the higher the division quality of the first community division proposal is.
[0092] In a more specific embodiment, the target attribute modularity of the target community is determined in step 206. include:
[0093] According to the first similarity sim(Z c ,Z) and the second target norm ‖Z‖ of the global embedding representation Z to determine the first ratio. The first ratio can be
[0094] Then, according to the first target norm ‖Z c ‖The weighted difference with the first ratio determines the modularity of the target attribute As shown in formula (1):
[0095]
[0096] Among them, ‖Z c‖ represents the target community embedding representation Z c The first objective norm of c The square of the 2-norm, that is, Z c With Z c λ represents the inner product between λ and λ. λ represents the preset weight. sim() is a similarity function, such as cosine similarity. ‖Z‖ represents the second target norm of the global embedding representation Z, for example, it can be the square of the 2-norm of Z, that is, the inner product between Z and Z.
[0097] In other embodiments, other specific calculation methods for determining the modularity of the target attribute may be set. For example, the first target norm ‖Z c ‖ minus the first similarity sim(Z c ,Z), to determine the target attribute modularity.
[0098] It can be understood that step 206 shows the process of calculating the attribute modularity of one of the communities. Each working node in the distributed system also needs to calculate the attribute modularity of other communities in the first community division proposal in a similar manner to step 206 and send it to the master node.
[0099] After the master node receives the attribute modularity of each community, next, in step 208, the master node determines the global attribute modularity Q according to the received attribute modularity of each community. a ; The global attribute modularity Q a Used to evaluate the division quality of the first community division proposal.
[0100] The master node can aggregate the attribute modularity of each community to determine the global attribute modularity Q a .
[0101] In one embodiment, the step 208 of determining the global attribute modularity includes: the master node determines the global attribute modularity Q according to the ratio of the sum of the attribute modularities of each community to the second target norm ‖Z‖ of the global embedding representation. a As shown in formula (2):
[0102]
[0103] Among them, C represents the first community partition proposal, and c∈C represents the community in the first community partition proposal.
[0104] The global attribute modularity Q a It is used to evaluate the partition quality of the first community partition proposal. Specifically, Q a When the value of is greater than 0, it means that the community division proposal is an effective division that can bring gains and can be used for subsequent updates of community division results. aWhen Q is greater than 0, a The larger the value of , the higher the division quality of the first community division proposal.
[0105] In other embodiments, the master node may also use other methods to aggregate the attribute modularity of each community to determine the global attribute modularity Q a For example, the master node can take the average of the attribute modularity of each community as the global attribute modularity Q a .
[0106] The above describes a community partition proposal (the first community partition proposal) generated for the master node, using distributed computing to calculate its global attribute modularity Q a In other implementations, the master node can also generate multiple community division proposals based on the existing community division results, and calculate the global attribute modularity Q of each community division proposal according to the above steps. a , and then select the global attribute module Q a The top-ranked community division proposals are used as the final community division proposals, and the existing community division results are updated.
[0107] The above describes the process of judging the partition quality of the community partition proposal based on the node attributes. In some possible implementations, the node attributes and the graph structure attributes can be combined to jointly judge the partition quality of the community partition proposal. In this implementation, the method further includes step 212 and step 214.
[0108] In step 212, the target working node determines the target structural modularity of the target community according to the relevant edge weights and the global weight W of the target community. And send it to the master node; the global weight is determined by aggregating the edge weights of each relationship edge in the full graph.
[0109] Target structural modularity The edge weight of the relationship edge associated with the target community is related. The relationship edge associated with the target community may be a relationship edge connecting graph nodes in the target community.
[0110] In one embodiment, step 212 determines the target structure modularity of the target community. Includes step 2122 and step 2124.
[0111] In step 2122, the target working node aggregates the edge weights of the internal relationship edges of the target community c to obtain the internal weight of the target community c. And aggregate the edge weights of the relationship edges connected to each graph node in the target community c to obtain the total weight of the target community
[0112] The internal relationship edge of the target community c can be a relationship edge in which the two connected graph nodes are both graph nodes in the target community c. The relationship edge connected to each graph node in the target community c can be a relationship edge in which at least one of the two connected graph nodes is a graph node in the target community c. The aggregation operation can be to calculate the sum of each weight or to calculate the average value of each weight, which is not limited here.
[0113] In a specific embodiment, the target working node obtains the internal weight of the target community c according to the sum of the edge weights of the internal relationship edges of the target community c. And according to the sum of the edge weights of the relationship edges connected to each graph node in the target community c, the total weight of the target community is obtained
[0114] Then, in step 2124, the target working node calculates the target community's internal weight. and the total weight And the global weight W, calculate the modularity of the target structure Make it positively correlated with the internal weight Negatively related to the total weight
[0115] Target structural modularity Positively correlated to the internal weight This means that the internal weight The larger the target structure, the higher the modularity The larger the value, the higher the partition quality of the first community partition proposal; the target structure modularity Negatively related to the total weight This means that the total weight The smaller the target structure modularity The larger it is, the higher the partition quality of the first community partition proposal is.
[0116] In a more specific embodiment, the step 2124 of calculating the modularity of the target structure include:
[0117] According to the internal weight of the target community The ratio between W and the global weight W determines the second ratio. The second ratio can be According to the total weight of the target community The ratio between W and the global weight W determines the third ratio. The third ratio can be
[0118] Then, the target structural modularity is determined according to the weighted difference between the second ratio and the square of the third ratio. As shown in formula (3):
[0119]
[0120] Among them, γ represents the preset weight.
[0121] In other embodiments, other specific calculation methods for determining the modularity of the target structure may be set, for example, by converting the internal weight Subtract the total weight To determine the modularity of the target structure.
[0122] It can be understood that step 212 shows the process of calculating the structural modularity of one of the communities. Each working node in the distributed system also needs to calculate the structural modularity of other communities in the first community division proposal in a manner similar to step 212 and send it to the master node.
[0123] In one embodiment, the global weight W is determined by the following steps:
[0124] Any working node determines the total node weight of the target graph node based on the sum of the edge weights of each relationship edge connected to any target graph node on the graph shard it holds, and sends it to the master node; the master node determines the global weight W based on the sum of the total node weights of each received node, and sends it to each working node.
[0125] Since the global weight W is not affected by the community division result, when the entire graph remains unchanged, the global weight W is always unchanged, so the above steps can be pre-executed before step 212. After receiving the global weight W, each working node can save it and use it in step 212 without recalculating it in each round of community division.
[0126] Then, in step 214, the master node receives the structural module degrees of each community. Determine the global structural modularity Q s The global structural modularity Q s For the global attribute modularity Q a Jointly evaluate the division quality of the first community division proposal.
[0127] The master node can aggregate the structural modularity of each community to determine the global structural modularity Q s .
[0128] In one embodiment, the master node determines the global structural modularity Q according to the sum of the structural modularity of each community. s , as shown in formula (4):
[0129]
[0130] Among them, C represents the first community partition proposal, and c∈C represents the community in the first community partition proposal.
[0131] Global structural modularity Q s For the global attribute modularity Q a The quality of the first community partition proposal is jointly evaluated. Specifically, the global structural modularity Q s and the global attribute modularity Q a Perform weighted summation to obtain the global total modularity Q, as shown in formula (5):
[0132] Q=βQ a +(1-β)Q s (5)
[0133] Among them, β is a preset parameter with a value between 0 and 1. The global total modularity Q is used to evaluate the partition quality of the first community partition proposal by comprehensively considering the graph node attributes and graph structure. Specifically, when the value of Q is greater than 0, it means that the community partition proposal is an effective partition that can bring gains and can be used for subsequent updates of community partition results. When Q is greater than 0, the larger the value of Q, the higher the partition quality of the first community partition proposal.
[0134] In one embodiment, the target community includes a target central graph node, and the target central graph node is located on a graph shard held by a target working node. In this embodiment, before determining the target structural modularity of the target community, the method further includes:
[0135] Step 210 : The target working node receives the edge weights of the relationship edges connecting the respective graph nodes from the working nodes holding the graph nodes of the target community.
[0136] The above describes a community partition proposal (the first community partition proposal) generated for the master node, using distributed computing to calculate its global structural modularity Q s , and then calculate the global total modularity Q to determine the specific steps of its division quality. In other implementations, the master node can also generate multiple community division proposals based on the existing community division results, and calculate the global structural modularity Q of each community division proposal according to the above steps. s , and then calculate the global total modularity Q, and then select the community division proposal with the highest global total modularity Q as the final community division proposal, and update the existing community division results.
[0137] The embodiments of this specification propose the use of distributed computing based on the attribute modularity of graph node attributes to improve the accuracy and rationality of community division, while also improving the computational efficiency of the modularity of large-scale relationship graphs and the timeliness of the output results. Furthermore, by using attribute modularity and structural modularity in combination, the contribution of graph node attributes and graph structure to community division can be considered simultaneously, and the weight of the relationship data between users and the attribute data of the users themselves in determining the modularity can be better balanced, thereby mining high-quality communities in different scenarios to further increase the rationality of community division.
[0138] According to an embodiment of another aspect, a distributed community discovery system is also provided. Figure 3 FIG. 1 is a schematic block diagram of a distributed community discovery system according to an embodiment, which can be deployed in any device, platform or device cluster with computing and processing capabilities. Figure 3 As shown, the distributed community discovery system includes a master node and several working nodes, each of which holds a graph slice of the full graph, wherein:
[0139] The master node is used to send the first community division proposal for the entire graph to each working node, which includes several communities;
[0140] Any target working node is used to determine the target community embedding representation based on the embedding representation of each graph node attribute in any target community on the held graph shard, and then determine the target attribute modularity of the target community, and send it to the master node; the target attribute modularity is negatively correlated to the first similarity between the target community embedding representation and the global embedding representation; the global embedding representation is determined based on the aggregation result of the embedding representation of each graph node attribute in the whole graph;
[0141] The master node is further used to determine the global attribute modularity according to the received attribute modularity of each community; the global attribute modularity is used to evaluate the partition quality of the first community partition proposal.
[0142] According to another aspect of the embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method described in any of the above embodiments.
[0143] According to yet another embodiment, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in any one of the above embodiments is implemented.
[0144] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0145] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0146] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0147] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0148] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A distributed community discovery method, applied to a distributed system, wherein the distributed system includes a master node and a plurality of worker nodes, each of which holds a graph shard of a full graph; the method includes: The master node sends a first community division proposal for the entire graph to each worker node, which includes several communities. Any target working node determines the target community embedding representation based on the embedding representation of each graph node attribute in any target community on the graph shard it holds, and then determines the target attribute modularity of the target community and sends it to the master node; The target attribute modularity is negatively correlated to the first similarity between the target community embedding representation and the global embedding representation; The global embedding representation is determined according to an aggregation result of the embedding representations of the attributes of each graph node in the whole graph; The master node determines the global attribute modularity based on the attribute modularity of each community received; The global attribute modularity is used to evaluate the division quality of the first community division proposal.
2. The method according to claim 1, wherein: The first community division proposal is an update proposal determined for an existing community division result, wherein there is at least one migration node, the community to which the migration node belongs in the first community division proposal is different from the community to which the migration node belongs in the existing community division result.
3. The method according to claim 2, further comprising: The master node moves at least one graph node in any community in the existing community division results to an adjacent community to obtain the first community division proposal.
4. The method according to claim 1, wherein: The target community includes a target central graph node, and the target central graph node is located on a graph shard held by a target working node; The step of determining the target community embedding representation comprises: The target worker node receives the embedded representation of each graph node attribute from each worker node holding the graph node of the target community; The target working node performs a first aggregation process on the embedded representations of the attributes of each graph node in the target community, thereby determining the embedded representation of the target community.
5. The method according to claim 1, wherein: The global embedding representation is determined by the following steps: Any working node determines the embedded representation of the shard based on the embedded representation of each graph node attribute on the graph shard it holds and sends it to the master node; The master node performs a second aggregation process based on the received embedding representations of each shard, thereby determining a global embedding representation and sending it to each working node.
6. The method according to claim 1, wherein: The target attribute modularity is positively correlated with the first target norm of the target community embedding representation.
7. The method according to claim 6, wherein: Determine the target attribute modularity of the target community, including: Determine a first ratio according to a ratio of the square of the first similarity to a second target norm of the global embedding representation; The target attribute modularity is determined according to a weighted difference between the first target norm and the first ratio.
8. The method according to claim 1, wherein: Determine global attribute modularity, including: The global attribute modularity is determined based on the ratio of the sum of the attribute modularity of each community to the second objective norm of the global embedding representation.
9. The method according to claim 1, further comprising: The target working node determines the target structural modularity of the target community according to the relevant edge weights and global weight of the target community, and sends it to the master node; The global weight is determined by aggregating the edge weights of each relationship edge in the whole graph; The master node determines the global structural modularity based on the structural modularity of each community received; The global structural modularity is used together with the global attribute modularity to evaluate the partition quality of the first community partition proposal.
10. The method according to claim 9, wherein: Determine the target structural modularity of the target community, including: The target working node aggregates the edge weights of the internal relationship edges of the target community to obtain the internal weight of the target community, and aggregates the edge weights of the relationship edges connected to each graph node in the target community to obtain the total weight of the target community; The target working node calculates the target structure modularity according to the internal weight and total weight of the target community and the global weight, so that the modularity is positively correlated with the internal weight and negatively correlated with the total weight.
11. The method according to claim 9, wherein: The target community includes a target central graph node, and the target central graph node is located on a graph shard held by a target working node; Before determining the target structural modularity of the target community, the method further includes: The target working node receives the edge weights of the relationship edges connecting each graph node from each working node holding the graph node of the target community.
12. The method according to claim 9, wherein: The global weight is determined by the following steps: Any working node determines the total node weight of the target graph node based on the sum of the edge weights of each relationship edge connected to any target graph node on the graph shard it holds, and sends it to the master node; The master node determines the global weight based on the sum of the total node weights of each node received and sends it to each working node.
13. The method according to claim 10, wherein: Calculating the modularity of the target structure includes: Determining a second ratio according to a ratio between an internal weight of the target community and a global weight; Determining a third ratio according to a ratio between the total weight of the target community and the global weight; The target structural modularity is determined according to a weighted difference between the second ratio and the square of the third ratio.
14. The method according to claim 1, wherein: The full graph is a user relationship graph, in which graph nodes represent users and relationship edges represent relationships between users; the community represents a group composed of several users.
15. A distributed community discovery system includes a master node and several worker nodes, each worker node holds a graph shard of the full graph, wherein: The master node is used to send the first community division proposal for the entire graph to each working node, which includes several communities; Any target working node is used to determine the target community embedding representation based on the embedding representation of each graph node attribute in any target community on the graph shard it holds, and then determine the target attribute modularity of the target community and send it to the master node; The target attribute modularity is negatively correlated to the first similarity between the target community embedding representation and the global embedding representation; The global embedding representation is determined according to an aggregation result of the embedding representations of the attributes of each graph node in the whole graph; The master node is also used to determine the global attribute modularity based on the attribute modularity of each community received; The global attribute modularity is used to evaluate the division quality of the first community division proposal.
16. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 14.
17. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 14 is implemented.