A Method for Obtaining Strongly Connected Components of Large-Scale Graph Data with High CPU Efficiency
By adding virtual node r to the directed graph and constructing memory sampling diagram A, combined with the depth-first search algorithm, the problem of low efficiency in obtaining strong connectivity components in the existing technology is solved, and efficient acquisition in a semi-external memory environment is achieved.
Patent Information
- Application Number
- CN202211138474.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-19
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-09-19
AI Technical Summary
The existing strongly connected component acquisition method requires exponential running time in a semi-external memory environment, resulting in low acquisition efficiency per unit time.
By adding virtual node r to the directed graph, the memory sampling graph A is constructed, and the edge scanning and shrinking is used using the depth-first search algorithm, the strong connected components on the directed graph G stored in the disk are gradually obtained, including the first and second shrinkage processes.
In a semi-external memory environment, strongly connected components acquisition at the cost of linear CPU is realized, improving the acquisition efficiency per unit time without consuming exponential running time.
Smart Images

Figure CN115481296B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graph computing for big data processing, and particularly relates to a method for obtaining strongly connected components of large-scale graph data with high CPU efficiency. Background Art
[0002] Obtaining strongly connected components has many applications in the field of graph theory. Given a directed graph G, all strongly connected components on G are returned, where a strongly connected component on G is the largest subgraph in which all nodes are mutually reachable. Existing in-memory algorithms for strongly connected components are very efficient, such as the Kosaraju-Sharir algorithm. However, with the continuous growth of the current graph data volume, it is often difficult to store G entirely in the memory of the computing device.
[0003] Currently, obtaining strongly connected components in the field of graph theory is mainly implemented using in-memory algorithms or external memory algorithms. External memory algorithms generally assume that the memory size of the computing device is M and the size of a disk block is B. Although the proposed external memory algorithms enable users to calculate the strongly connected components of G on computing devices with any memory size, precisely because of this, external memory algorithms usually involve very complex operations and it is difficult to efficiently return the results to users. The external memory strongly connected component calculation method proposed by Cosgaya-Lozano et al. can return the results to users quickly in some cases, but this method is found to fall into an infinite loop in certain cases. The LS algorithm proposed by Laura et al. only requires an integer array of length 2n to return all strongly connected components of G, but this method takes time to process each edge on G. In-memory methods need to import all graph data into memory, and external memory methods need to ensure that the algorithm can still run when the memory space is very small. Therefore, semi-external memory methods for strongly connected components have been actively discussed in recent years.
[0004] The semi-external memory strongly connected component method sets a lower bound on the size of the memory environment, which requires that the memory space of the computing device can store at least one spanning tree for graph G. However, the current semi-external memory strongly connected component calculation algorithms consume exponential I / O. Therefore, it is still very difficult to calculate the strongly connected components of graph data in a semi-external memory environment. The semi-external memory depth-first search method EdgeByBatch (EB-DFS) proposed by Sibeyn et al. can be used to solve the semi-external memory strongly connected component problem. However, this method requires a large amount of I / O in practice and is difficult to meet the user's requirements for query efficiency. Zhang et al. designed a series of semi-external memory strongly connected component calculation methods for this problem. Later, all these methods were found to still require a large amount of I / O and CPU calculations. Wan et al. proposed an efficient method for the semi-external memory strongly connected component calculation method, named EP-SCC. Although EP-SCC has been tested on some graph models and the experimental results show that this method has a certain degree of efficiency, the EP-SCC method still cannot obtain the strongly connected components in a semi-external memory environment with a linear CPU cost. Therefore, the current methods for obtaining strongly connected components still have the problem of consuming exponential running time, resulting in low efficiency of obtaining strongly connected components per unit time. Summary of the Invention
[0005] The object of the present invention is to solve the problem that the existing methods for obtaining strongly connected components consume exponential running time, resulting in low efficiency of obtaining strongly connected components per unit time, and to propose a method for efficiently obtaining strongly connected components of large-scale graph data by CPU.
[0006] The specific process of a method for efficiently obtaining strongly connected components of large-scale graph data by CPU is as follows:
[0007] Step 1: Obtain a directed graph stored on disk, add a virtual node r to the directed graph, and use the directed graph G with the added virtual node to obtain a memory sampling graph A and the set Ei of edges in G;
[0008] The memory sampling graph A includes all node sets on G;
[0009] Each node u in the memory sampling graph A has two attribute values: u.H and u.L;
[0010] Among them, both u.H and u.L are integers;
[0011] All other nodes in the directed graph G with the added virtual node are connected to r by a directed edge;
[0012] Step 2: Use A and Ei obtained in Step 1 to obtain all strongly connected components on the directed graph G stored on disk, including the following steps:
[0013] Step 2.1: Initialize the set of edges E0 in the memory sampling graphs A and G, and set the variable i = 0;
[0014] Step 2.2: Determine whether Ei is empty. If Ei is empty, directly obtain all strongly connected components on G based on A, and then end. If Ei is not empty, execute Step 2.3;
[0015] Step 2.3: Use to scan the set Ei, and then determine whether all edges in Ei have been scanned. If all edges in Ei have been scanned, execute Step 2.4. If not all edges in Ei have been scanned, execute Step 2.5;
[0016] Step 2.4: Determine whether the relative order between nodes has changed during the scanning of Ei. If there is a change, set i = i + 1, and then execute Step 2.2. If the relative order between nodes has not changed during the scanning of Ei, obtain all strongly connected components on G based on A, and then end;
[0017] Step 2.5: Scan any edge e(u, v) in G, and obtain the corresponding nodes s and t of u and v in A. Determine whether s and t are reachable in A or s is equal to t. If they are reachable or s = t, execute Step 2.3. If they are not reachable and s ≠ t, add (s, t) to A, and then execute Step 2.6;
[0018] where e and v are two different nodes;
[0019] Step 2.6: Determine whether A can be further enlarged. If A can still be further enlarged, execute Step 2.3. If A cannot be further enlarged, perform a contraction on A. If A can be further enlarged, save a spanning tree with node r in A in the memory, and store the remaining edge data except the spanning tree to the disk Ei+1, and then execute Step 2.3 until Ei is empty or the relative order of nodes in A does not change during the scanning of Ei, and obtain all strongly connected components on G.
[0020] The contraction of A includes: the first contraction and the second contraction.
[0021] The specific process of a method for obtaining strongly connected components of large-scale graph data with high CPU efficiency is as follows:
[0022] Step 1: Obtain a directed graph stored on the disk, add a virtual node r to the directed graph, and use the directed graph G with the added virtual node to obtain a memory sampling graph A;
[0023] The memory sampling graph A includes all node sets on G;
[0024] Each node u in the memory sampling graph A has two attribute values: u.H and u.L;
[0025] where u.H and u.L are integers;
[0026] All other nodes in the directed graph G with virtual nodes added are connected to r by a directed edge;
[0027] Step 2: Use A obtained in Step 1 to obtain all strongly connected components on the directed graph stored on disk, including the following steps:
[0028] Step 2-1: Initialize A;
[0029] Step 2-2: Scan the directed graph G and determine whether all edges in G have been scanned. If all edges in G have been scanned, obtain all strongly connected components on G based on A; if there are edges in G that have not been scanned, execute Step 2-3;
[0030] Step 2-3: Scan any edge e=(u, v) on G and obtain the corresponding nodes s and t of nodes u and v in A;
[0031] Step 2-4: Determine whether there is a reachability relationship between s and t on A or s=t. If there is a reachability relationship between s and t on A or s=t, execute Step 2-2; if there is no reachability relationship between s and t on A and s≠t, execute Step 2-5;
[0032] Step 2-5: After adding (s, t) to A, determine whether A can continue to be enlarged. If A can continue to be enlarged, execute Step 2-2; if A cannot continue to be enlarged, contract A to obtain the contracted A, and then execute Step 2-2 until all edges in G have been scanned to obtain all strongly connected components on G.
[0033] The beneficial effects of the present invention are:
[0034] The present invention proposes a method for a linear CPU in a semi-external memory environment to obtain strongly connected components of a directed graph, which ensures the strongly connected components in the external memory environment while ensuring that the computational CPU cost is within a linear range. At the same time, the present invention does not require exponential running time consumption, improving the acquisition efficiency of strongly connected components per unit time. Description of the Drawings
[0035] Figure 1 is a flowchart of the first specific embodiment of the present invention;
[0036] Figure 2 is a flowchart of the first sub-function of the present invention;
[0037] Figure 3 is a flowchart of the second sub-function;
[0038] Figure 4 is a flowchart of the second specific embodiment of the present invention;
[0039] Figure 5 It is the flowchart of sub-function three. Detailed implementation manners
[0040] Detailed implementation manner one: As Figure 1 shown, the specific process of a method for obtaining strongly connected components of large-scale graph data with high CPU efficiency in this implementation manner is as follows:
[0041] Step 1: Obtain a directed graph stored on disk, add a virtual node r to the directed graph, and use the directed graph G with the added virtual node to obtain a memory sampling graph A and a set Ei of edges in G:
[0042] All node sets on G are included in the memory sampling graph A;
[0043] Each node u in the memory sampling graph A has two attribute values: u.H and u.L;
[0044] All other nodes in the directed graph G with the added virtual node are connected to r by a directed edge;
[0045] Step 2: Use A and Ei obtained in Step 1 to obtain all strongly connected components on the directed graph G stored on disk:
[0046] Step 2.1: Initialize the memory sampling graph A, the set E0 of edges in G, and the variable i = 0;
[0047] where i is the number of edges in the set;
[0048] Step 2.2: Determine whether Ei is empty. If Ei is empty, directly obtain all strongly connected components on G based on A, and then end; if Ei is not empty, execute Step 2.3;
[0049] Step 2.3: Scan the set Ei, and then determine whether all edges in Ei have been scanned. If all edges in Ei have been scanned, execute Step 2.4; if not all edges in Ei have been scanned, execute Step 2.5;
[0050] Step 2.4: Determine whether the relative order between nodes has changed during the scanning of Ei. If it has changed, set i = i + 1, and then execute Step 2.2; if the relative order between nodes has not changed during the scanning of Ei, obtain all strongly connected components on G based on A, and then end;
[0051] Step 2.5: Scan any edge e(u, v) in G, and obtain the corresponding nodes s and t of u and v in A. Determine whether s and t are reachable in A or s is equal to t. If they are reachable or s = t, execute Step 2.3; if they are not reachable and s ≠ t, add (s, t) to A, and then execute Step 2.6;
[0052] Among them, e and v are two different nodes;
[0053] Whether the reachability is determined by the following method;
[0054] If s is not an ancestor of t on T, then there is no reachability relationship between s and t; if s is an ancestor of t on T, then there is a reachability relationship between s and t;
[0055] Step 26: Determine whether A can continue to be increased. If A can continue to be increased, then execute Step 23. If A cannot continue to be increased, then contract A; if A can continue to be increased, then save a spanning tree T in A that has a node r in memory, and store the remaining edge data except the spanning tree to disk Ei+1, and then execute Step 23 until Ei is empty or the relative order of the nodes scanned in A does not change, and all strongly connected components on G are obtained.
[0056] The contraction of A includes: the first contraction and the second contraction.
[0057] Specific implementation method 2; As Figure 2 shown, the first contraction is implemented by using the depth-first search algorithm, that is, sub-function 1, and specifically includes the following steps:
[0058] S1. Initialize the tree T in A that has a node r, the empty edge set E, and the set S;
[0059] The set S only includes one node r;
[0060] S2. Determine whether S is empty. If S is empty, then obtain the graph composed of T and E, that is, the result of the first contraction; if S is not empty, then execute S3;
[0061] S3. Assume that the top element of the stack of S is u1, and determine whether u1 has been visited. If u1 has been visited, then execute S4; if u1 has not been visited, then execute S5;
[0062] S4. Remove u1 from S, and then determine whether u1 is a node on T. If u1 is a node on T, then u1.H = the number of nodes that have been visited currently, and then execute S2; if u1.H is not a node on T, then execute S2;
[0063] S5. Set u1.L = the number of nodes that have been visited currently, then mark u1 as a visited node, and when u1 is not equal to r, add (u.H, u1) to T;
[0064] S6. Scan all adjacent nodes v of u1 on A, then determine whether v has been visited. If v has not been visited, add v to S and set v.H = u1; if v has been visited and v is a node on T, record v; if v has been visited and v.H is not a node on T, store v in set V.
[0065] S7. After all adjacent nodes of node u1 have been scanned, select a node y with the smallest value of.L from all the recorded nodes v, and contract the nodes from y to u1 into one node on T.
[0066] S8. After all adjacent nodes of node u1 have been scanned, if V is not empty, add the edge (x, u1) to E, then execute S2. If V is empty, directly execute S2 until S2 outputs the graph composed of T and E, which is the result of the first contraction.
[0067] Where x is any element in V.
[0068] In this embodiment, the first sub-function uses depth-first search as the basic framework, that is, it uses a stack structure S. In the sub-function, the initialization of S includes an element r. The symbols T and E in the first sub-function are used to represent the intermediate results of the two calculation processes of the sub-function.
[0069] Specific Embodiment 3: As Figure 3 shown, the second contraction is implemented using the second sub-function, including the following steps:
[0070] 1) Obtain the depth-first search order of each node in the first contraction process, that is, the order in which the first sub-function visits the nodes, and obtain the generated search tree T using the depth-first search order.
[0071] 2) Determine whether all nodes in A after the first contraction have been visited. If not all nodes have been visited, execute 3); if all nodes have been visited, directly output the result, that is, A after the second contraction.
[0072] 3) Obtain the node u' with the largest depth-first search order among the currently unvisited nodes, then determine whether there is an ancestor w of u' on T that has not been visited. If there is an unvisited w and w is not equal to r, set u = w, and then execute 4); if there is no unvisited w or w is equal to r, execute 4).
[0073] 4) Mark u' as having been visited, then start traversing all nodes in A from u', obtain all reachable nodes x of u' on A, and mark x as having been visited.
[0074] 5) Retain all the edges of A on T or the edges on the search tree during the access process, remove the remaining edges on A, and then execute 2) until the second contraction of A is obtained.
[0075] Embodiment 4: As Figure 4 shown, a method for obtaining strongly connected components of large-scale graph data with high CPU efficiency includes the following steps:
[0076] Step 1: Obtain the directed graph stored on disk, add a virtual node r to the directed graph, and use the directed graph G with the added virtual node to obtain the in-memory sampling graph A;
[0077] The in-memory sampling graph A includes all the node sets on G;
[0078] Each node u in the in-memory sampling graph A has two attribute values: u.H and u.L;
[0079] Among them, u.H and u.L are integers;
[0080] All other nodes in the directed graph G with the added virtual node are connected to r by a directed edge;
[0081] Step 2: Use A obtained in Step 1 to obtain all the strongly connected components on the directed graph stored on disk, including the following steps:
[0082] Step 2.1: Initialize A;
[0083] Step 2.2: Scan the directed graph G, and determine whether all the edges in G have been scanned. If all the edges in G have been scanned, obtain all the strongly connected components on G based on A; if there are edges in G that have not been scanned, execute Step 2.3;
[0084] Step 2.3: Scan any edge e=(u, v) on G, and obtain the corresponding nodes s and t of node u and node v in A;
[0085] Step 2.4: Determine whether there is a reachability relationship between s and t on A or s=t. If there is a reachability relationship between s and t on A or s=t, execute Step 2.2; if there is no reachability relationship between s and t on A and s≠t, execute Step 2.5;
[0086] Step 2.5: After adding (s, t) to A, determine whether A can continue to be enlarged. If A can continue to be enlarged, execute Step 2.2; if A cannot continue to be enlarged, contract A to obtain the contracted A, and then execute Step 2.2 until all the edges in G have been scanned to obtain all the strongly connected components on G.
[0087] The reachability is determined in the following way;
[0088] If s is not an ancestor of t on T, then there is no reachability relationship between s and t; if s is an ancestor of t on T, then there is a reachability relationship between s and t.
[0089] Specific Embodiment 5: As Figure 5 shown, if A cannot be increased further, A is shrunk, and the shrunk A is implemented using Sub-function 3, which specifically includes the following steps:
[0090] S1. Initialize the tree T with node r in A, the empty edge set E, and the set S;
[0091] Among them, the initialized set S includes an element r;
[0092] S2. Determine whether S is empty. If S is empty, obtain the graph composed of T and E, which is the shrunk A; if S is not empty, execute S3;
[0093] S3. Assume that the top element of S is u1, and determine whether u1 has been visited. If u1 has been visited, execute S4; if u1 has not been visited, execute S5;
[0094] S4. Remove u1 from S, and then determine whether u1 is a node on T. If u1 is a node on T, then u1.H = the number of nodes that have been visited, and then execute S2; if u1.H is not a node on T, then execute S2;
[0095] S5. u1.L = the number of nodes that have been visited, then mark u1 as a visited node, and when u1 is not equal to r, add (u.H, u1) to T;
[0096] S6. Scan all adjacent nodes v of u1 in A, and then determine whether v has been visited. If v has not been visited, add v to S, and at the same time v.H = u1; if v has been visited and v is a node on T, then record v; if v has been visited and v.H is not a node on T, then store v in the set V;
[0097] S7. After scanning all adjacent nodes of node u1, select a node y with the smallest attribute value.L from all the recorded nodes v, and contract the nodes from y to u1 into one node on T;
[0098] S8. After scanning all adjacent nodes of node u1, if V is not empty, add the edge (x, u1) to E, and then execute S2. If V is empty, directly execute S2 until S2 outputs the graph composed of T and E, which is the shrunk A;
[0099] Among them, x is any element in V.
[0100] The content of the present invention is not limited to the content of the above embodiments, and the combination of one or several specific embodiments can also achieve the purpose of the invention.
Claims
1. A method for obtaining strongly connected components of large-scale graph data with high CPU efficiency, characterized in that The specific process of the method is as follows: Step 1: Obtain the directed graph stored on the disk, add a virtual node r to the directed graph, and use the directed graph G with the added virtual node to obtain the in-memory sampling graph A and the set Ei of edges in G. The in-memory sampling graph A includes all the node sets on G. Each node u in the in-memory sampling graph A has two attribute values: u.H and u.L. Among them, both u.H and u.L are integers. All nodes in the directed graph G with the added virtual node are connected to r by a directed edge. Step 2: Use A and Ei obtained in Step 1 to obtain all the strongly connected components on the directed graph G stored on the disk, including the following steps: Step 2.1: Initialize the in-memory sampling graph A, the set E0 of edges in G, and the variable i = 0. Step 2.2: Determine whether Ei is empty. If Ei is empty, directly obtain all the strongly connected components on G based on A, and then end. If Ei is not empty, execute Step 2.
3. Step 2.3: Scan the set Ei, and then determine whether all the edges in Ei have been scanned. If all the edges in Ei have been scanned, execute Step 2.
4. If not all the edges in Ei have been scanned, execute Step 2.
5. Step 2.4: Determine whether the relative order between nodes changes during the scanning of Ei. If it changes, let i = i + 1, and then execute Step 2.
2. If the relative order between nodes does not change during the scanning of Ei, obtain all the strongly connected components on G based on A, and then end. Step 2.5: Scan any edge e(u, v) in G, and obtain the corresponding nodes s and t of u and v in A. Determine whether s and t are reachable in A or s is equal to t. If they are reachable or s = t, execute Step 2.
3. If they are not reachable and s ≠ t, add (s, t) to A, and then execute Step 2.
6. Among them, v and u are two different nodes. Step 2.6: Determine whether A can continue to be enlarged. If A can still be enlarged, execute Step 2.
3. If A cannot be enlarged, contract A. If A can be enlarged, save a spanning tree T with the node r in A in memory, and store the remaining edge data except the spanning tree on the disk Ei+1, and then execute Step 2.3 until Ei is empty or the relative order of nodes in A does not change during the scanning of Ei, and obtain all the strongly connected components on G. The contraction of A includes: the first contraction and the second contraction. The second contraction includes the following steps: 1) Obtain the depth-first search order of each node during the first contraction process, and use the depth-first search order to obtain the generated search tree T. 2) Determine whether all the nodes in A after the first contraction have been visited. If not all the nodes have been visited, execute 3). If all its nodes have been visited, directly output the result, that is, A after the second contraction. 3) Obtain the node u' with the largest depth - first search order among the currently unvisited nodes, and then determine whether there is an ancestor w of u' in T that has not been visited. If there exists an unvisited w and w is not equal to r, then let u = w, and then execute 4); if there is no unvisited w or w is equal to r, then execute 4); 4) Mark u' as visited, and then start traversing all nodes in A from u', obtain all reachable nodes x of u' in A, and mark x as visited; 5) Retain all the edges in A that are on T, remove the remaining edges in A, and then execute 2) until the second - contracted A is obtained.
2. The method for obtaining strongly connected components of large-scale graph data with high CPU efficiency according to claim 1, wherein: The determination of reachability is made in the following way: If s is not an ancestor of t, then there is no reachability relationship between s and t; if s is an ancestor of t, then there is a reachability relationship between s and t.
3. The method for obtaining strongly connected components of large-scale graph data with high CPU efficiency according to claim 2, wherein: The first contraction is implemented using a depth - first search algorithm, specifically as follows: S1. Initialize the tree T in A that has the node r, the empty edge set E, and the set S; S2. Determine whether S is empty. If S is empty, then obtain the graph composed of T and E, which is the result of the first contraction; if S is not empty, then execute S3; S3. Assume that the top - element of the stack of S is u1, and determine whether u1 has been visited. If u1 has been visited, then execute S4; if u1 has not been visited, then execute S5; S4. Remove u1 from S, and then determine whether u1 is a node on T. If u1 is a node on T, then u.H = the current number of visited nodes, and then execute S2; if u1.H is not a node on T, then execute S2; S5. Let u1.L = the current number of visited nodes, then mark u1 as a visited node, and in the case where u1 is not equal to r, add (u.H, u) to T; S6. Scan all adjacent nodes v of u1 in A, and then determine whether v has been visited. If v has not been visited, then add v to S, and at the same time v.H = u1; if v has been visited and v is a node on T, then record v; if v has been visited and v.H is not a node on T, then store v in the set V; S7. After scanning all adjacent nodes of the node u1, select a node y with the smallest attribute value.L from all the recorded nodes v, and contract the nodes from y to u1 into one node on T; S8. After scanning all adjacent nodes of the node u1, if V is not empty, add the edge (x, u1) to E, and then execute S2. If V is empty, then directly execute S2 until S2 outputs the graph composed of T and E, which is the result of the first contraction; where x is any element in V.
4. The method for obtaining strongly connected components of large-scale graph data with high CPU efficiency according to claim 3, wherein: After the set S is initialized, S only includes one node r.