A distributed graph segmentation method and device based on neighbor node clustering

By segmenting and communication optimization of graphs in distributed computing clusters, the problems of high complexity and poor scalability of existing algorithms when processing giant graphs are solved, and efficient graph segmentation and load balancing are achieved.

CN118170950BActive Publication Date: 2025-05-23ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311781141.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-05-23
Estimated Expiration
2043-12-22

AI Technical Summary

Technical Problem

The existing distributed graph segmentation algorithm has high complexity, poor scalability when processing giant graphs, and is difficult to adapt to the situation of different node weights.

Method used

A distributed graph segmentation method based on neighbor node clustering is proposed. By segmenting the nodes and edges of the graph in a distributed computing cluster, load balancing and cutting edge minimization are achieved using All-to-all ensemble communication and a stand-alone LDG algorithm.

Benefits of technology

It realizes efficient segmentation of giant graphs in a distributed environment, reduces communication overhead between server hosts, and has the advantages of simplicity, small time overhead and strong scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118170950B_ABST
    Figure CN118170950B_ABST
Patent Text Reader

Abstract

The present invention proposes a distributed graph partitioning method and device based on neighbor node clustering. Construct an undirected graph and a specified number of partitions, divide the point set into several subsets, so that each node is located in and only in one of the subsets; meet the load balancing restriction, the partition size is as equal as possible; minimize the number of cut edges, that is, the total number of edges with two ends in different partitions is as small as possible. The present invention is based on the LDG algorithm for graph partitioning problem: adopt the idea of ​​neighbor node clustering, assign each node to the partition with the most neighbor nodes, and reduce the number of cut edges by incorporating closely related nodes into the same partition; apply a linear penalty factor to larger partitions, so as to meet the load balancing restriction. On the basis of the single-machine LDG algorithm, the present invention designs and implements a set of efficient, easy-to-implement distributed parallel computing solutions suitable for any number of partitions k, which can complete the partitioning of giant graph structures in a relatively short running time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of distributed storage and computing of graph data structures, and in particular to a distributed graph segmentation method and device based on neighbor node clustering. Background Art

[0002] Graph data structures play a vital role in the field of big data storage and computing. Social networks, transportation and logistics networks, and epidemic spread in real life can all be modeled as a graph G(V,E), where the node v∈V represents each social network user, logistics distribution point, or disease infected person, and the edge e=(u,v)∈E represents the relationship between nodes, such as social network attention, transportation roads, or epidemic spread chains. As the scale of the graph structure continues to expand, it will reach an order of magnitude that a single server host can no longer accommodate (10 9 Even higher), the distributed storage of graphs has become an important factor affecting the performance of graph databases and graph computing engines, so the research on graph segmentation algorithms came into being: distributing the nodes and edges of the graph structure to multiple hosts in the distributed storage architecture, and achieving load balancing as much as possible and minimizing the distributed communication overhead during the graph computing process.

[0003] Graph partitioning algorithms include two types: point partitioning and edge partitioning. The former constitutes the mainstream of related research and applications. Divide the point set V of the graph into a specified number of k subsets and minimize the number of edges across partitions under the premise of load balancing. Related studies have proposed a variety of algorithm paradigms such as iterative heuristic optimization, neighbor node clustering, and multi-layer graph partitioning. Among them, neighbor node clustering represented by LDG, that is, giving priority to selecting the partition with the most neighbors, has been favored by many graph computing solutions because of its simple implementation, low computational overhead and good effect. However, existing research (including algorithm design and experimental design) lacks sufficient involvement in distributed graph partitioning algorithms for giant graphs, and existing research on distributed algorithms either has complicated implementation and huge algorithm complexity constants; or has limitations in algorithm design and is only applicable to k=2. i , or the special case where the node weights are all 1. In view of the above problems, the present invention transforms the LDG algorithm into a distributed parallel one, provides a detailed step design for the distributed communication strategy in the algorithm process, and can be easily extended to the case where the number of partitions k and the number of processors P are any positive integers respectively, and the nodes can have different weights. Summary of the invention

[0004] The purpose of the present invention is to address the deficiencies of the prior art and propose a distributed graph segmentation method and device based on neighbor node clustering. The present invention assigns the information of each node of the graph to a suitable storage node, ensures the uniformity of storage and computing overhead through load balancing, and reduces the communication overhead between server hosts by minimizing edge cutting.

[0005] The objective of the present invention is achieved through the following technical solutions: In a first aspect, the present invention proposes a distributed graph segmentation method based on neighbor node clustering, the method comprising the following steps:

[0006] (1) There are P hosts in the distributed computing cluster. All computing nodes in the distributed computing cluster form an undirected graph structure G(V,E), where V represents the set of computing nodes and E represents the set of edges between computing nodes.

[0007] (2) Divide the point set V and edge set E into P subsets respectively, and use the subsets of the edge set and the weights of all computing nodes in the subsets of the point set as the initial input of the corresponding host;

[0008] (3) All computing nodes are connected through all-to-all collective communication between hosts The weight c(u) is moved to the i-th host; all hosts that meet the starting point Move the edge (u,v) of the i-th host to obtain the preprocessed edge set. Indicated as a label in r i-1 ~r i -1 All computing nodes within the range; r i Indicates the node number demarcation point of the i-th host computer;

[0009] (4) Within the i-th host, define For all edges whose endpoints are located at the jth host, For the above All the starting nodes that have appeared in, i≠j, are convenient for determining the target host of the communication between nodes;

[0010] (5) For each computing node v∈V, let π(v)←RandInt(1,k), that is, randomly select a partition number with equal probability between 1 and k;

[0011] (6) Perform several iterations, and update the partition result π(v) in each iteration; get the i-th host to output all internal nodes The final partition label is π(u).

[0012] Furthermore, in step (3), the host where each computing node is located is indexed according to the obtained processing results. For each computing node v∈V, it is only necessary to use the label of v to locate which host in the computing cluster uses it as an internal node.

[0013] Furthermore, in step (3), the specific steps are as follows:

[0014] (3.1) Within each host, each undirected edge Split into two edges (u,v) and (v,u);

[0015] (3.2) Execute any distributed sorting algorithm that guarantees load balancing after sorting on all the split edges in the lexicographic order of (u,v), so that after the algorithm ends, the edge set held by the first host is The second host holds the 2|E| / P edges with the smallest lexicographic order. Hold the next 2|E| / P edges... and so on until the last host

[0016] (3.3) Calculate the demarcation point corresponding to each host That is, each node u is assigned to the host with the smallest label when it appears as the starting point of the edge;

[0017] (3.4) Order For the label in r i-1 ~r i -1 All nodes within the range of all nodes through the All-to-all collective communication between hosts The weight c(u) is moved to the i-th host;

[0018] (3.5) Input the strategy label A or B. If maxdeg(v)>2|E| / P, the strategy label is forced to be changed to B; if the strategy label is B, jump to step 3.7; otherwise, jump to step 3.6;

[0019] (3.6) All the starting points that satisfy Move the edge (u, v) of the i-th host to obtain the preprocessed edge set End the preprocessing step;

[0020] (3.7) For each partition order

[0021] (3.8) Order Represents all "cross-partition" nodes u: they are internal nodes of another j-th host However, there is an edge starting from u in the current i-th host. Complete all After the calculation, it is notified to the jth host through the All-to-all collective communication operator, ending the preprocessing step.

[0022] Furthermore, in step (6), each round of iteration process is as follows:

[0023] (6.1) Synchronous partition result: The i-th host will communicate through All-to-all collection The partition results of all nodes in the cluster are sent to the jth host;

[0024] (6.2) Single-machine LDG algorithm: Each host executes the single-machine LDG algorithm with k and λ as parameters for all internal nodes, and updates all internal nodes The partition π(u) where .

[0025] Furthermore, the steps of the single-machine LDG algorithm are as follows:

[0026] 1) Calculate the total weight of all internal nodes And the upper limit L of each partition's capacity i =(1+λ)C i / k;

[0027] 2) Record the margin L of each partition i,j , the initial margin of the jth partition is R i,j ←L i ;

[0028] 3) Traverse all internal nodes: For each node Select a partition label for it That is, after taking the margin of each partition as the penalty factor, the partition with the most neighbor nodes, and then let R i,π(u) ←R i,π(u) -c(u) Reduce the margin of the selected partition.

[0029] In a second aspect, the present invention also provides a distributed graph segmentation device based on neighbor node clustering, comprising a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, the distributed graph segmentation method based on neighbor node clustering is implemented.

[0030] In a third aspect, the present invention further provides a computer-readable storage medium having a program stored thereon, and when the program is executed by a processor, the distributed graph segmentation method based on neighbor node clustering is implemented.

[0031] Beneficial effects of the present invention: Compared with the partitioning algorithms based on multi-layer graph contraction such as METIS, the algorithm based on neighbor node clustering proposed in the present invention has the advantages of simple distributed implementation, low running time overhead, and strong scalability. Experimental results on multiple large graph structures show that if the number of cut edges is used as a criterion, the algorithm of the present invention can achieve a segmentation effect similar to that of METIS in a shorter time. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0033] Figure 1 Flowchart of the dist-LDG algorithm of the present invention.

[0034] Figure 2 The following is an example diagram of the internal node dependency between hosts.

[0035] Figure 3 Schematic diagram of the preprocessing step (edge ​​storage adopts strategy A).

[0036] Figure 4 Schematic diagram of the preprocessing step (edge ​​storage adopts strategy B).

[0037] Figure 5 It is a structural diagram of a distributed graph segmentation device based on neighbor node clustering of the present invention. DETAILED DESCRIPTION

[0038] The specific implementation modes of the present invention are further described in detail below with reference to the accompanying drawings.

[0039] The present invention provides a distributed graph segmentation method based on neighbor node clustering. It is mainly used in social networks, which have the most typical power-law graph characteristics. Specifically, a distributed graph segmentation algorithm dist-LDG based on the LDG algorithm is adopted, and the input is a social network undirected graph structure G(V,E) and other parameters of the algorithm. The nodes in the undirected graph represent users in the social network, and the edges represent the relationship between users, such as the attention of the social network. A small number of nodes representing "important users" have a huge number of edges connected to them, which means that "important users" have higher influence and are easy to spread social network events. The weight c(v) of each node v∈V is>0, and the sum of the weights of all nodes in the graph is recorded as C. In the present invention, the nodes and edges of the graph structure are distributed to multiple hosts in the distributed storage architecture, and load balancing and minimization of distributed communication overhead in the graph calculation process are achieved as much as possible. The algorithm of the present invention calculates the partition number π(v)∈{1,…,k} of each node v∈V in a distributed computing cluster with P hosts as the output of the algorithm. The algorithm meets the following requirements:

[0040] (1) Load balancing: the total weight of each partition is no more than 1+λ times the mean, that is,

[0041] (2) Minimize the number of cut edges. The algorithm reduces the total number of cut edges as much as possible. The predicate function It is equal to 1 if the statement x is true, and 0 otherwise.

[0042] The algorithm flow of dist-LDG is as follows Figure 1 As shown, the following steps are included:

[0043] 1. Distributed input: divide the edge set E of the relationship between users in the undirected graph into P subsets in an arbitrary way in As the initial input of the i-th host; divide the user point set V of the undirected graph into P subsets in an arbitrary way in The weights of all nodes in are used as the initial input of the i-th host; for each host, the total number of nodes N, the number of partitions k, the imbalance factor λ, and the number of iterations T are input;

[0044] 2. Check the legality of the input to ensure that there are no duplicate edges, no self-loops, the number of nodes N ≥ max{k, P}, the weight of each user node c(v)>0, the degree of each user node deg(v)>0, and the number of each edge node is in the range of 1,…,N;

[0045] 3. Preprocess the input. Within each host, convert each undirected edge representing the relationship between users Split into two edges (u,v) and (v,u); calculate the node number dividing point 1 = r 0 <r 1 <… <r P =N+1, the “internal node set” V of the i-th host i (in) For the label in r i-1 ~r i All user nodes within the range of -1, all nodes u∈V i (in) The weight c(u) is moved to the i-th host; all hosts that meet the starting point Move the edge (u,v) to the i-th host and get the preprocessed edge set

[0046] 4. Maintain the dependencies between hosts. Within the i-th host, define For all edges whose endpoints are located at the jth host, For the above All starting nodes that have appeared in, i≠j;

[0047] 5. Random initial solution: for each node v∈V, let π(v)←RandInt(1,k), that is, randomly select a partition number with equal probability between 1 and k;

[0048] 6. Perform T rounds of iterations, and update the partition result π(v) in each round of iteration;

[0049] 7. Output results: the i-th host outputs all internal nodes u∈V i (in) The final partition label is π(u).

[0050] Step 3 is used to facilitate indexing of the host where each node is located: for each node v∈V in the whole graph, only the label of v itself can be used to locate which host in the cluster uses it as an internal node; for each internal node u∈V i (in) , the present invention moves all the neighboring edges of u to the host where u is located, so that in the subsequent step 6 of the single-machine LDG algorithm, all the neighboring edges of u can be accessed without any distributed communication operation. The fourth step is to determine the target host for communication between nodes. Figure 2 As shown, for an internal node u∈V of the i-th host i (in) , and any edge whose endpoint is adjacent to another host After preprocessing, the jth host has its reverse edge in memory Before traversing the partition status of all neighbor nodes of v in the subsequent step 6, it is necessary to synchronize π(u) from the i-th host to the j-th host through all-to-all collective communication, and vice versa.

[0051] In the preprocessing process of step 3 of dist-LDG, the cutoff point r 0 …r P There are many setting strategies. The naive approach is to evenly divide the node numbers 1 to N, but this scheme is prone to load imbalance during the graph partitioning algorithm for power-law graph structures with uneven node degree distribution (a few nodes have a huge number of edges connected to them). For this reason, the following optimized preprocessing steps are proposed:

[0052] 3.1In each host, each undirected edge Split into two edges (u,v) and (v,u);

[0053] 3.2 Execute any distributed sorting algorithm that guarantees load balancing after sorting on all the split edges in the lexicographic order of (u,v), so that after the algorithm ends, the edge set held by the first host The second host holds the 2|E| / P edges with the smallest lexicographic order. Hold the next 2|E| / P edges... and so on until the last host

[0054] 3.3 Calculate the demarcation point corresponding to each host That is, each node u is assigned to the host with the smallest label when it appears as the starting point of the edge;

[0055] 3.4 Let V i (in) For the label in r i-1 ~r i All nodes within the range of -1, all nodes u∈V i (in) The weight c(u) is moved to the i-th host;

[0056] 3.5 Input the strategy label A or B. If maxdeg(v)>2|E| / P, the strategy label is forced to be changed to B; if the strategy label is B, jump to step 3.7; otherwise, jump to step 3.6;

[0057] 3.6 All starting points u∈V i (in) Move the edge (u,v) to the i-th host and get the preprocessed edge set End the preprocessing step;

[0058] 3.7 For each partition order

[0059] 3.8 Order Represents all "cross-partition" nodes u: they are internal nodes V of another j-th host j (in) However, there is an edge starting from u in the current host i. i (in) Complete all After the calculation, it is notified to the jth host through the All-to-all collective communication operator, ending the preprocessing step.

[0060] The above strategy A can ensure that each |V after preprocessing is i (in) When |>0, there is an upper limit on the load of each host. by Figure 3 As shown in the figure, the red, blue, and green colors correspond to the three hosts in the computing cluster when P=3. The top is the distribution of the lower side after preprocessing is completed. The lower right corner shows the sub-images in each host after preprocessing. , where the colored points represent V i (in) Internal points, solid line edges indicate that both ends are located at V i(in) The dotted line indicates that only the starting end is located at V i (in) Inside.

[0061] The above strategy B may appear |V in the extreme case of maxdeg(v)>2|E| / P i (in) |=0, but the upper limit of each host load is further reduced to |V|+2|E| / P. Obviously, by Figure 4 As an example, the top shows the distribution of the bottom after preprocessing. The golden dot in the lower right diagram indicates All the rest All are empty.

[0062] For step 6 above, the process of each iteration is as follows:

[0063] 6.1 Synchronize partition results. The i-th host will communicate with the All-to-all collection The partition results of all nodes in the cluster are sent to the jth host;

[0064] 6.2 Single-machine LDG algorithm, each host executes the single-machine LDG algorithm for all internal nodes with k,λ as parameters, and updates all internal nodes The partition π(u) where .

[0065] Based on the specific implementation of the above preprocessing operation, step 6.2 (taking the i-th host as an example) is expanded as follows:

[0066] 6.2.1 Calculate the total weight of all internal nodes And the upper limit of each partition's capacity L i =(1+λ)C i / k;

[0067] 6.2.2 Record the margin L of each partition i,j , the initial margin of the jth partition is R i,j ←L i ;

[0068] 6.2.3 Set the untraversed node set

[0069] 6.2.4 If End this step, otherwise jump to step 6.2.5;

[0070] 6.2.5 Take any node u∈U i , let U i ←U i-{u}; If step 3 uses strategy B and there is a j-th host such that Then jump to step 6.2.7, otherwise jump to 6.2.4;

[0071] 6.2.6 Local Computation Partition Label That is, after taking the margin of each partition as the penalty factor, the partition with the most neighbor nodes, and then let R i,π(u) ←R i,π(u) -c(u) reduces the margin of the selected partition and jumps to step 6.2.4;

[0072] 6.2.7 Locally Calculate the Total Weight of Each Partition

[0073] 6.2.8 For each like Then send a calculation request to the jth host, which will calculate the total weight of each partition locally. Then, the Send / Receive collection communication operator is used to return to the current i-th host, and then

[0074] 6.2.9 Calculating the partition index π(u)←argmax 1≤r≤k [L i,r -c(v)]s r , then R i,π(u) ←R i,π(u) -c(u), jump to step 6.2.4.

[0075] In the above steps, the margin R i,j The restrictions ensure that each host can meet Therefore, after all P hosts are accumulated, the global load balance is also guaranteed. In step 6.2.3, each node u is preferentially assigned to the partition with the most neighboring user nodes, which minimizes the number of cut edges in u's neighboring edges and promotes the clustering within each partition. The present invention is applicable to distributed parallel computing solutions with any number of partitions k, and can complete the partitioning of giant graph structures in a shorter running time.

[0076] Corresponding to the aforementioned embodiment of a distributed graph segmentation method based on neighbor node clustering, the present invention also provides an embodiment of a distributed graph segmentation device based on neighbor node clustering.

[0077] See also Figure 5A distributed graph segmentation device based on neighbor node clustering provided in an embodiment of the present invention includes a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, it is used to implement a distributed graph segmentation method based on neighbor node clustering in the above embodiment.

[0078] An embodiment of a distributed graph segmentation device based on neighbor node clustering provided by the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the internal memory for execution. From a hardware perspective, if Figure 5 As shown, it is a hardware structure diagram of a distributed graph segmentation device based on neighbor node clustering provided by the present invention, in which any device with data processing capability is located, except Figure 5 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in which the apparatus in the embodiments is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.

[0079] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.

[0080] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.

[0081] An embodiment of the present invention further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, a distributed graph segmentation method based on neighbor node clustering in the above embodiment is implemented.

[0082] The computer-readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device of any device with data processing capability, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capability. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.

[0083] The above embodiments are used to illustrate the present invention rather than to limit the present invention. Any modification and change made to the present invention within the spirit of the present invention and the protection scope of the claims shall fall within the protection scope of the present invention.

Claims

1. A distributed graph segmentation method based on neighbor node clustering, It is characterized in that The method comprises the following steps: (1) There are P hosts in the distributed computing cluster. All computing nodes in the distributed computing cluster form an undirected graph structure G(V,E), where V represents the set of computing nodes and E represents the set of edges between computing nodes. (2) Divide the point set V and edge set E into P subsets respectively, and use the weights of all computing nodes in the edge set subsets and point set subsets after division as the initial input of the corresponding host; (3) All computing nodes u∈V are connected through all-to-all collective communication between hosts. i (in) Move the weight c(u) to the i-th host; move all hosts that satisfy the starting point u∈V i (in) The edge (u, v) is moved to the i-th host to obtain the preprocessed edge set, V i (in) Indicated as a label in r i-1 ~r i -1 All computing nodes within the range; r i Indicates the node number demarcation point of the i-th host computer. The specific steps are as follows: (3.1) Within each host, each undirected edge Split into two edges (u,v) and (v,u); (3.2) Execute any distributed sorting algorithm that guarantees load balancing after sorting on all the split edges in the lexicographic order of (u,v), so that after the algorithm ends, the edge set held by the first host is The second host holds the 2|E| / P edges with the smallest lexicographic order. Hold the next 2|E| / P edges... and so on until the last host (3.3) Calculate the demarcation point corresponding to each host That is, each node u is assigned to the host with the smallest label when it appears as the starting point of the edge; (3.4) Let V i (in) For the label in r i-1 ~r i All nodes within the range of -1, all nodes u∈V i (in) The weight c(u) is moved to the i-th host; (3.5) Input the strategy label A or B; if maxdeg(v)>2|E| / P, the strategy label is forced to be changed to B; if the strategy label is B, jump to step 3.7; otherwise, jump to step 3.6; (3.6) Move all edges (u, v) where the starting point u ∈ V i (in) to the i-th host to obtain the preprocessed edge set End the preprocessing step; (3.7) For each partition order (3.8) Order i>j represents all "cross-partition" nodes u: they are internal nodes V of another j-th host. j (in) , but there is an edge starting from u in the current i-th host; from the stored Complete all After the calculation, it is notified to the jth host through the All-to-all collective communication operator, and the preprocessing step ends; (4) Within the i-th host, define For all edges whose endpoints are located at the jth host, For the above All the starting nodes that have appeared in, i≠j, are convenient for determining the target host of the communication between nodes; (5) For each computing node v∈V, let π(v)←RandInt(1,k), that is, randomly select a partition number with equal probability between 1 and k; (6) Perform several iterations, and update the partition result π(v) in each iteration; get the output of all internal nodes u∈V by the i-th host. i (in) The final partition number is π(u); each round of iteration process is as follows: (6.1) Synchronous partition result: The i-th host will communicate with the The partition results of all nodes in the cluster are sent to the jth host; (6.2) Single-machine LDG algorithm: Each host executes the single-machine LDG algorithm with k and λ as parameters for all internal nodes, and updates all internal nodes u∈V i (in) The partition π(u) where .

2. A distributed graph segmentation method based on neighbor node clustering according to claim 1, It is characterized in that In step (3), the host where each computing node is located is indexed according to the processing results. For each computing node v∈V, it is only necessary to use the label of v to locate which host in the computing cluster uses it as an internal node.

3. A distributed graph segmentation device based on neighbor node clustering, comprising a memory and one or more processors, wherein the memory stores executable code, It is characterized in that When the processor executes the executable code, a distributed graph segmentation method based on neighbor node clustering as described in any one of claims 1-2 is implemented.

4. A computer-readable storage medium having a program stored thereon, It is characterized in that When the program is executed by a processor, a distributed graph segmentation method based on neighbor node clustering as described in any one of claims 1-2 is implemented.

Citation Information

Patent Citations

  • Distributed user clustering method for social network

    CN112633388A

  • Partitioning social networks

    US20060015588A1