Query method for elastic expansion graph loaded on demand
By combining on-demand loading and elastic expansion, the existing graph query system's problems of insufficient resource prediction, large loading overhead and high preprocessing overhead are solved, efficient graph data processing and dynamic expansion are achieved, and system performance is improved.
Patent Information
- Application Number
- CN202510430546.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
AI Technical Summary
The existing distributed graph query system has problems such as insufficient resource prediction, large loading overhead and high preprocessing overhead when processing large-scale graph data, which is difficult to adapt to dynamically changing business needs. The existing elastic expansion graph query system has performance bottlenecks in communication and re-division overhead.
Using the elastic expansion graph query method that loads on demand, by building an improved CSR index file, managing nodes control data loading, and computing nodes load data required by local active graph algorithms on demand, avoiding redundant data loading and preprocessing, and dynamically loading vertex and edge data using graph topology, reducing cross-node edge data, and achieving elastic expansion of storage and computing separation.
It effectively reduces disk and network I/O overhead, avoids preprocessing overhead, improves data locality and resource utilization, achieves good dynamic elastic expansion capabilities, and significantly improves the performance of the graph query system.
Smart Images

Figure CN120353971A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graph query, and more specifically, to an elastic expansion graph query method with on-demand loading. Background Art
[0002] A graph is an abstract data structure composed of vertices and edges, which can intuitively express the relationships between entities. Vertices represent objects in the real world (such as users in a social network, stations in a transportation network), and edges can represent the relationships between objects (such as friendship relationships, transportation routes).
[0003] Graph algorithms analyze and mine graph data to reveal information such as relationship patterns, community structures, and key nodes hidden in the data. The single-source shortest path algorithm (SSSP) can calculate the shortest distance between two nodes in an urban road network graph, which is of great significance to fields such as navigation and physics; breadth-first search (BFS) is a traversal algorithm that systematically explores a graph by level, and is commonly used to measure network connectivity, find minimum spanning trees, explore links for web crawlers, etc.; the PageRank algorithm can determine the importance of web pages by analyzing the link relationships between web pages.
[0004] To better execute graph algorithms, the academic community has proposed graph query systems specifically designed for processing and analyzing graph data. However, with the exponential growth of Internet data volume, stand-alone graph query systems face bottlenecks in memory limitations and computing efficiency, and it is difficult to process graph data at the TB level. Therefore, the academic community and the industrial community have launched many distributed graph query systems to meet the needs of large-scale graph data processing.
[0005] Existing distributed graph query systems rely on a distributed architecture to disperse huge graph data across multiple machines, overcoming the resource and computing power limitations of stand-alone graph query systems. They generally adopt a "vertex-centric" or "edge-centric" computing model, and accelerate the speed of large-scale graph analysis by processing multiple nodes and edges simultaneously. At the same time, preprocessing techniques (graph partitioning, vertex reordering) are used to further accelerate the execution speed of graph algorithms. However, in order to pursue the highest possible computing efficiency, they generally adopt an architecture design with integrated storage and computing (binding computing and storage resources), ignoring the scalability of the system. This design pattern with highly coupled computing and storage results in limited scalability of the system when facing different scales of workloads, and it is difficult to adapt to dynamic business requirements.
[0006] In summary, although existing graph query systems all show very good performance when processing large-scale graph data, these graph query systems are all based on an integrated storage and computing system architecture, and there are the following problems:
[0007] 1. Resource prediction. Since these systems cannot scale elastically, it is necessary to predict the computing resource requirements in advance. Different graph sizes and workloads require different amounts of computing resources. For example, larger graphs require more computing resources. If the computing resources are insufficient, the computation needs to be stopped and the distributed cluster needs to be reconfigured.
[0008] 2. Loading overhead. Before executing graph algorithms, these graph query systems need to read all graph data. However, for algorithms such as SSSP and BFS, only the data related to the vertices reachable from the source vertex is required to complete the computation. Loading all graph data will result in additional I / O overhead and cause resource waste.
[0009] 3. Preprocessing overhead. Existing graph query systems rely on expensive preprocessing techniques (such as graph partitioning) to accelerate graph queries, which leads to the cold start of graph query tasks. In most cases, the total time of preprocessing and computation will exceed the time required to directly execute the computation without preprocessing.
[0010] In recent years, with the rapid development of cloud computing, the compute-storage separation architecture has become the mainstream system architecture used in cloud platforms, and at the same time, it has provided a new idea for realizing the elastic expansion of graph query systems, which can effectively solve the first problem mentioned above. This architecture decouples and manages computing resources and storage resources separately, bringing flexible elastic expansion capabilities to the system. In a compute-storage separation graph query system, the storage layer can be independently expanded to cope with the growing graph data scale, while the computing layer can dynamically adjust resource allocation according to the computational complexity and the volume of concurrent requests to achieve on-demand scaling. This design effectively improves resource utilization and reduces computing costs. Currently, there have been some explorations on elastic expansion of graph query systems in the academic community, mainly Graphless and ElGA, which use the compute-storage separation architecture to solve the first problem mentioned above: the scalability of graph query systems.
[0011] "Graphless: Toward Serverless Graph Processing" published in the conference ISPDC 2019 constructs the first graph query system based on the serverless architecture, aiming to make graph queries more accessible to ordinary users. This system integrates Amazon's AWS cloud services, loads graph data from Amazon's S3 object storage, uses AWS's Lambda functions to perform distributed computing, and relies on the Redis database as a storage service to maintain vertex states and message passing. Graphless simply combines serverless computing with graph queries and is elastically expanded by the cloud platform.
[0012] "ElGA: Elastic and Scalable Dynamic Graph Analysis", published in the conference SC 2021, designed an elastic and scalable dynamic graph query system based on the separation of storage and computing. It adopted a shared-nothing architecture and a consistent hashing method, which can dynamically adjust computing resources according to the growth of the graph or computing requirements. The design of the shared-nothing architecture enables each computing node to run independently and communicate only through message passing, facilitating the addition or removal of computing nodes. Consistent hashing allows the system to determine the global location of edges and maintain it even when nodes are added or removed. When nodes are expanded, only a small portion of data needs to be moved, minimizing data migration during the expansion operation.
[0013] "Graphless: Toward Serverless Graph Processing", published in the conference ISPDC 2019, directly combined serverless functions with the graph query system without optimizing serverless graph queries. Although it has relatively good performance in lightweight algorithms such as BFS, when executing computationally intensive algorithms such as PageRank, due to its reliance on remote memory services to maintain state persistence and communication between worker nodes, this results in excessive message passing overhead. Therefore, Graphless still performs worse than traditional graph query systems when dealing with communication-intensive workloads.
[0014] The consistent hashing adopted by "ElGA: Elastic and Scalable Dynamic Graph Analysis" published in the conference SC 2021 provides good elasticity, but its performance in graph partitioning usually cannot be compared with specialized graph partitioning algorithms (such as METIS), resulting in a large repartitioning overhead during expansion.
[0015] Graphless and ElGA have made preliminary explorations on the elastic expansion of the graph query system, but there are still problems with system performance: Graphless uses serverless functions for computing and relies on remote storage services to maintain vertex state and communication between worker nodes. Graphless needs to communicate frequently, resulting in a large amount of communication overhead. ElGA uses consistent hashing to implement graph partitioning, facilitating the location of the global position of edges for elastic expansion, but the graph partitioning efficiency of consistent hashing is low, resulting in additional preprocessing overhead. In addition, these systems cannot avoid high preprocessing overhead, and there is no data loading optimization for algorithms such as SSSP and BFS. The full data loading results in additional loading overhead. Summary of the Invention
[0016] Aiming at the deficiencies of the prior art, the object of the present invention is to propose an elastic expansion graph query method with on-demand loading, including:
[0017] Step 1: For a directed acyclic graph, construct an improved CSR index file. The improved CSR index file includes an index file and an edge file. The edge file contains multiple lines of data, and each line of data includes the vertex pointed to in the directed acyclic graph, or each line of data contains the vertex pointed to in the directed acyclic graph and the weight of the edge between the pointing vertex and the pointed vertex. The index file includes the pointing vertex and the starting line number of the pointing vertex in the edge file;
[0018] Step 2: Input the storage path of the index file and the storage path of the edge file at a pre-configured storage node. The storage node reads the index file from the local disk according to the storage path of the index file, constructs an index array for each vertex according to the index file. The index array of the vertex includes the starting line number of the vertex in the edge file. The storage node sends the number of vertices to the computing node and sends the storage location of the vertex to the management node;
[0019] Step 3: A pre-configured management node receives the locally active graph algorithm to be executed input by the user and the source point in the directed acyclic graph. The locally active graph algorithm is an algorithm that can complete the execution task only with partial data of the directed acyclic graph; take the source point as the current active vertex. The management node sends the data loading task of the current active vertex to the storage node and sends the current active vertex and the locally active graph algorithm to a pre-configured current computing node;
[0020] Step 4: After receiving the data loading task of the current active vertex, the storage node reads the edge file from the local disk according to the storage path of the edge file; according to the index array of the current active vertex, obtains the edge data of multiple pointed vertices corresponding to the current active vertex in the edge file, that is, the edge data of multiple target vertices, and sends the edge data to the current computing node, and pre-reads the edge data of the pointed vertices corresponding to the target vertices;
[0021] Step 5: The current computing node receives all the current active vertices and the locally active graph algorithm, initializes the state values of all vertices in the directed acyclic graph according to the algorithm type of the locally active graph algorithm; receives the edge data corresponding to the current active vertex sent by the storage node, and updates the state values of all target vertices according to the locally active graph algorithm, takes the target vertices as the new current active vertices, and obtains multiple current active vertices;
[0022] Step 6: Determine the size of the edge data corresponding to all the current active vertices, and obtain the memory capacity of the current computing node. According to the size of the edge data corresponding to all the current active vertices and the memory capacity of the obtained current computing node, determine the target computing node;
[0023] Step 7: The target computing node sends all current active vertices to the management node, and at the same time determines whether the edge data of the target vertex corresponding to the current active vertex has been stored in the target computing node among all current active vertices. If there is edge data of the target vertex corresponding to the current active vertex stored in the target computing node among all current active vertices, the status value of the target vertex corresponding to the current active vertex is updated according to the local active graph algorithm. If there is no edge data of the target vertex corresponding to the current active vertex stored in the target computing node among all current active vertices, the target computing node remains in a waiting state;
[0024] Step 8: The management node receives all current active vertices. The management node sends the data loading task of the current active vertices to the storage node. After receiving the data loading task of the current active vertices, the storage node obtains the edge data of multiple pointed-to vertices of the current active vertices read in advance, that is, the edge data of multiple target vertices, and sends the edge data of multiple target vertices to the target computing node, and then similarly reads the edge data of the pointed-to vertices of the target vertex in advance;
[0025] Step 9: After receiving the edge data of multiple target vertices, the target computing node determines that there is no target vertex that has been newly updated in Step 7 among all target vertices, and updates its status value according to the local active graph algorithm;
[0026] Step 10: Determine whether the target vertex meets the termination condition. If the target vertex meets the termination condition, store the final status values of all vertices from the source point to the target vertex. If the target vertex does not meet the termination condition, use all updated target vertices as the current active vertices and return to execute Step 6.
[0027] Optionally, in Steps 3 and 8, when the management node sends the data loading task of the current active vertices to the storage node, it includes:
[0028] The management node determines the storage node for storing the current active vertices and sends the current active vertices to the corresponding storage node.
[0029] Optionally, in Step 4, obtaining the edge data of multiple pointed-to vertices corresponding to the current active vertex in the edge file includes:
[0030] Obtain the index array of the current active vertex, and according to the starting line number of the current active vertex in the edge file in the index array, obtain the edge data of multiple pointed-to vertices corresponding to the current active vertex in the edge file.
[0031] Optionally, when the edge file in step 4 contains the vertices pointed to in the directed acyclic graph, the edge data of the target vertex only contains the target vertex pointed to by the current active vertex. When the edge file contains the vertices pointed to in the directed acyclic graph and the weights of the edges between the pointing vertices and the pointed vertices, the edge data of the target vertex contains the target vertex pointed to by the current active vertex and the weight of the edge from the current active vertex to the target vertex.
[0032] Optionally, the local active graph algorithm includes the shortest path algorithm SSSP, the breadth-first search algorithm BFS, the personalized page rank algorithm PPR, the penalty hit probability algorithm PHP, and the single-source widest path algorithm SSWP.
[0033] Optionally, step 6 specifically includes:
[0034] Determine the sizes of the edge data corresponding to all current active vertices, and obtain the memory capacity of the current computing node. Determine whether the memory capacity of the current computing node is less than the sizes of the edge data corresponding to all current active vertices. If the memory capacity of the current computing node is less than the sizes of the edge data corresponding to all current active vertices, the current computing node sends an expansion signal to the management node. The management node receives the expansion signal, expands new computing nodes, and the management node sends the routing information of the already loaded vertices, that is, the computing nodes where the already loaded vertices are located, to the new computing nodes. The current computing node sends the final status values of all vertices from the source point to the current active vertex stored locally to the new computing nodes, uses the new computing nodes as the target computing nodes, and executes step 7. If the memory capacity of the current computing node is greater than or equal to the sizes of the edge data corresponding to all current active vertices, use the current computing node as the target computing node and execute step 7.
[0035] Optionally, the termination conditions in step 10 include that the amount of state change is less than the termination threshold or the state no longer changes.
[0036] The beneficial effects produced by adopting the above technical solutions are as follows:
[0037] For local active graph algorithms such as SSSP and BFS, the present invention proposes an elastic extended graph query method with on-demand loading, that is, only loading the data participating in the calculation. Specifically, the management node receives the local active graph algorithm to be executed input by the user and the source points in the directed acyclic graph, and then updates the state values of the source points and the state values of some vertices in the directed acyclic graph, thereby avoiding the loading of redundant data when executing the local active graph algorithm and reducing the disk and network I / O overhead. At the same time, because on-demand loading loads the edge data of vertices in real time during the execution of the graph algorithm and does not require preprocessing the data set through graph partitioning, the preprocessing overhead of traditional graph query systems is avoided. At the same time, on-demand loading dynamically loads vertices and edges to the computing nodes according to the graph topology, so as to ensure that most of the closely related vertices are located on the same computing node. This topology loading method enhances the data locality of the graph query task and reduces the edge data across nodes, still achieving an effect similar to traditional graph partitioning. Existing elastic extended graph systems need to perform a large amount of data migration after expansion. The present invention combines on-demand loading with elastic expansion and does not use preprocessing technologies such as graph partitioning, avoiding the re-partitioning overhead during the elastic expansion process. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 FIG. is a system architecture diagram of a computer integrated system for implementing an elastic extended graph query method with on-demand loading in an embodiment of the present invention;
[0039] Figure 2 FIG. is a flow schematic diagram of an elastic extended graph query method with on-demand loading in an embodiment of the present invention;
[0040] Figure 3 FIG. is an execution time diagram of a local active graph algorithm in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0041] The following will further describe in detail the specific embodiments of the present invention in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0042] Aiming at the problems existing in the prior art, the purpose of the present invention is to design an elastic extended graph query method with on-demand loading in a distributed graph query application to reduce the redundant data loading overhead when executing algorithms such as SSSP and BFS, reduce the preprocessing overhead without affecting the system performance, and in addition, have good dynamic elastic expansion ability during the execution of computing tasks to ensure that the system can adapt to different demands of various workloads for computing resources. Thus, the distributed graph query system can meet the constantly changing actual application scenarios.
[0043] The on-demand loading mechanism of the present invention targets local active graph algorithms executed on directed graphs, such as the single-source shortest path algorithm (SSSP), breadth-first search algorithm (BFS), personalized page rank algorithm (PPR), penalty hit probability algorithm (PHP), and single-source widest path algorithm (SSWP). Local active means that only partial graph data is required to complete the computing task. Such algorithms only need to calculate all vertices reachable from the source point and do not require the complete graph data.
[0044] Specifically, the present invention provides an on-demand loading elastic extended graph query method. The present invention is implemented based on a computer integrated system. The architecture diagram of the computer integrated system is as Figure 1 , and the computer integrated system consists of a management node, a storage node, and a computing node. The management node is used to manage the cluster, mainly allocate data transmission tasks to the storage node, and control the expansion of the computing node. The storage node is responsible for transmitting data to the computing node according to the transmission tasks allocated by the management node. The computing node is responsible for executing computing tasks.
[0045] First, introduce the calculation process of the local active graph algorithm. The following active vertices refer to the vertices whose own states change after receiving messages from other vertices, except for the first-round active vertices. The first-round active vertices are determined according to the states of each vertex after initialization:
[0046] First, perform initialization. The local active graph algorithm initializes the vertex state values according to specific rules.
[0047] Then, select a source point s as the starting position of the algorithm's calculation. The source point s becomes active. At this time, the source point is the "first-round active vertex". Traverse all out-edge neighbor vertices x, y, z of s, execute the calculation of the local active graph algorithm, and update the state values of x, y, z.
[0048] The source point becomes inactive, and these updated out-edge neighbor vertices x, y, z are used as the "second-round active vertices" to perform the same calculation. Subsequently, update the state values of the out-edge neighbor vertices u, v, w of x, y, z.
[0049] x, y, z become inactive, and the vertices u, v, w whose state values have been updated are used as the "third-round active vertices" to perform the calculation and update the state values of the out-edge neighbor vertices of these vertices.
[0050] Repeat the above steps until all vertices reachable from the source point are visited or the change amount of the state values of all vertices is less than a certain threshold, and the calculation terminates.
[0051] The present invention combines such local active graph algorithms with on-demand loading to implement graph algorithms in the form of on-demand loading. First, it reads the CSR index file to generate an index array "index" for locating edge data. Then, the management node receives an execution command from the user, such as executing a local active graph algorithm with the source vertex being "v". The management node notifies the storage node to read the edge data of the source vertex, and then the storage node sends the data to the computing node. After the computing node receives the data required for calculation, it starts to calculate the active vertices in the first round. After the calculation is completed, the vertices in this round become inactive, and the vertices whose distance information is updated during the calculation become the active vertices in the next round. If the edge data of these vertices has not been loaded into the computing node, the above steps are repeated to load the data and then execute the calculation until the task terminates.
[0052] Specifically, in combination with Figure 2 , the following steps may be included:
[0053] Step 1: For a directed acyclic graph, construct an improved CSR index file. The improved CSR index file includes an index file and an edge file. The edge file contains multiple lines of data, and each line of data includes the vertex pointed to in the directed acyclic graph, or each line of data contains the vertex pointed to in the directed acyclic graph and the weight of the edge between the pointing vertex and the pointed vertex. The index file includes the pointing vertex and the starting line number of the pointing vertex in the edge file.
[0054] Among them, CSR (Compressed Sparse Row) is a commonly used sparse matrix storage format in graph queries. The CSR format consists of three arrays: value array (values): stores the weights or values of all edges in the graph. Column index array (index): stores the target vertices of each edge. Row offset array (offsets): stores the starting positions of the edges of each source vertex in the first two arrays. If you want to find all the edges of vertex "i", it is searched through the "offsets" array, and its out-edges are in the range of offsets[i] to offsets[i + 1] - 1 in the "index" array, that is, index[offsets[i]] to index[offsets[i + 1] - 1].
[0055] Among them, the improved CSR index file combines the column index and row offset of CSR into the index file. The present invention quickly locates the edge data that needs to be loaded in the graph data file through the index file.
[0056] Step 2: Input the storage paths of the index file and the edge file into the pre-configured storage node. The storage node reads the index file from the local disk according to the storage path of the index file, and constructs an index array for each vertex based on the index file. The index array of a vertex includes the starting line number of the vertex in the edge file. For example, the index array of a vertex can be expressed as index[s]=0, which means the starting position of vertex s in the edge file is 0. The storage node sends the number of vertices to the computing node to facilitate the initialization of the computing node. Further, the maximum vertex id among all vertices can also be sent to the computing node. For example, for four vertices 0, 1, 2, 3, the maximum vertex id is 3. The storage position of the vertex is sent to the management node to facilitate the management node to allocate data loading tasks. Since the index file is much smaller than the graph data file, the preprocessing time is significantly reduced.
[0057] Step 3: The pre-configured management node receives the local active graph algorithm to be executed input by the user, and the source vertex in the directed acyclic graph. The local active graph algorithm is an algorithm that can complete the execution task only with part of the data in the directed acyclic graph. Take the source vertex as the current active vertex. The management node sends the data loading task of the current active vertex to the storage node. Specifically, the management node determines the storage node storing the current active vertex, sends the current active vertex to the corresponding storage node, and sends the current active vertex and the local active graph algorithm to the pre-configured current computing node.
[0058] Among them, the local active graph algorithms include the shortest path algorithm SSSP, the breadth-first search algorithm BFS, the personalized page rank algorithm PPR, the penalty hit probability algorithm PHP, and the single-source widest path algorithm SSWP.
[0059] Step 4: After receiving the data loading task of the current active vertex, the storage node reads the edge file from the local disk according to the storage path of the edge file. According to the index array of the current active vertex, the edge data of multiple pointed-to vertices corresponding to the current active vertex, that is, the edge data of multiple target vertices, is obtained in the edge file. Specifically, the index array of the current active vertex is obtained, and according to the starting line number of the current active vertex in the edge file in the index array, the edge data of multiple pointed-to vertices corresponding to the current active vertex is obtained in the edge file. The edge data is sent to the current computing node, and the edge data of the pointed-to vertices corresponding to the target vertex is pre-read.
[0060] In the case where the edge file contains the vertices pointed to in the directed acyclic graph, the edge data of the target vertex only contains the target vertex pointed to by the current active vertex. For example, the edge files of the PPR, PHP, and BFS algorithms; in the case where the edge file contains the vertices pointed to in the directed acyclic graph and the weights of the edges between the pointing vertices and the pointed vertices, the target vertex edge data contains the target vertex pointed to by the current active vertex and the weight of the edge from the current active vertex to the target vertex. For example, the edge files of the SSSP and SSWP algorithms.
[0061] Step 5: The current computing node receives all the current active vertices and the local active graph algorithm, initializes the state values of all vertices in the directed acyclic graph according to the algorithm type of the local active graph algorithm; receives the edge data corresponding to the current active vertices sent by the storage node, updates the state values of all target vertices according to the local active graph algorithm, and takes the target vertices as the new current active vertices to obtain multiple current active vertices.
[0062] Among them, for the update of the state values of the target vertices, different algorithms have different update methods. For example, BFS is equivalent to the SSSP algorithm that treats all edge weights as 1, so weights are not required; for PPR and PHP, the state values are updated by the out-degree. The out-degree is the number of edges going out from a vertex. For example, if there are 5 out-edges, the message it sends to the destination vertex is its own state value divided by 5, and then the destination vertex accumulates the messages sent by all the vertices pointing to it.
[0063] Step 6: Determine the size of the edge data corresponding to all the current active vertices, and obtain the memory capacity of the current computing node. According to the size of the edge data corresponding to all the current active vertices and the obtained memory capacity of the current computing node, determine the target computing node.
[0064] Determine the size of the edge data corresponding to all the current active vertices, and obtain the memory capacity of the current computing node. Judge whether the memory capacity of the current computing node is less than the size of the edge data corresponding to all the current active vertices. In the case where the memory capacity of the current computing node is less than the size of the edge data corresponding to all the current active vertices, the current computing node sends an expansion signal to the management node. The management node receives the expansion signal, expands new computing nodes, and the management node sends the routing information of the vertices that have been loaded, that is, the computing nodes where the vertices that have been loaded are located, to the new computing nodes. The current computing node sends the final state values of all the vertices from the source point to the current active vertices stored locally to the new computing nodes, takes the new computing nodes as the target computing nodes, and executes Step 7. In the case where the memory capacity of the current computing node is greater than or equal to the size of the edge data corresponding to all the current active vertices, take the current computing node as the target computing node and execute Step 7.
[0065] Step 7: The target computing node sends all current active vertices to the management node, and at the same time determines whether the edge data of the target vertex corresponding to the current active vertex has been stored in the target computing node among all current active vertices. In the case where the edge data of the target vertex corresponding to the current active vertex has been stored in the target computing node among all current active vertices, the status value of the target vertex corresponding to the current active vertex is updated according to the local active graph algorithm. In the case where the edge data of the target vertex corresponding to the current active vertex has not been stored in the target computing node among all current active vertices, the target computing node remains in a waiting state;
[0066] Step 8: The management node receives all current active vertices. The management node sends the data loading task of the current active vertices to the storage node. Specifically, the management node determines the storage node storing the current active vertices and sends the current active vertices to the corresponding storage nodes. After receiving the data loading task of the current active vertices, the storage node obtains the edge data of multiple pointed-to vertices of the current active vertices read in advance, that is, the edge data of multiple target vertices, and sends the edge data of the multiple target vertices to the target computing node. Then, by the same token, it reads the edge data of the pointed-to vertices of the target vertex in advance;
[0067] Step 9: After receiving the edge data of multiple target vertices, the target computing node determines the target vertices that have not been updated just in Step 7 among all target vertices and updates their status values according to the local active graph algorithm;
[0068] That is to say, the present invention adopts a semi-blocking method. When reading the edge data of the target vertex corresponding to the current active vertex, it determines whether the computing node contains the edge data of the target vertex corresponding to the current active vertex. If the computing node stores the edge data of the target vertex corresponding to the current active vertex, the status value of the target vertex corresponding to the current active vertex is updated, thereby reducing the waiting time of the computing node.
[0069] Although the semi-blocking calculation has maximally utilized the computing resources, if all active vertices have not been loaded into the computing node, the computing node can only be in an idle state. To further reduce the data waiting time of the computing node and improve the reading efficiency, the storage node adopts a data prefetching strategy based on graph topology, that is, in Steps 4 and 8, it prefetches the edge data of the pointed-to vertices corresponding to the target vertex in advance. Because the storage node prefetches this data during idle time, the loading waiting time of the computing node can be reduced when performing the next round of updates.
[0070] Step 10: Determine whether the target vertex meets the termination condition. When the target vertex meets the termination condition, store the final state values of all vertices between the source vertex and the target vertex. Specifically, it can be written into a local file. When the target vertex does not meet the termination condition, return all the updated target vertices as the current active vertices and execute Step 6.
[0071] Among them, the termination condition includes that the state change amount is less than the termination threshold or the state no longer changes.
[0072] The present invention proposes an on-demand loading elastic expansion graph query method. On-demand loading means precise acquisition of data. Compared with traditional graph query methods, when executing local active graph algorithms such as SSSP, BFS, PPR, SSWP, and PHP, the present invention avoids data redundancy caused by loading the entire graph, reduces the loading overhead, and improves data utilization. As the graph query task is executed, the storage node continuously transmits data to the computing node. The present invention can dynamically adjust the number of computing nodes according to the resource usage situation, realize real-time node expansion, and improve the scalability of the present invention. Finally, on-demand loading avoids the preprocessing overhead of existing graph query methods, but still has a good graph partitioning effect. Since vertices and edges are dynamically loaded according to the graph topology, it is ensured that tightly connected vertices are located on the same computing node. This topology loading method enhances the data locality of the computing node and reduces the edges across nodes. In addition, the semi-blocking calculation of the computing node and the data prefetch of the storage node overlap data loading and calculation, improving resource utilization.
[0073] The present invention designs an on-demand loading elastic expansion graph query method, which ensures the scalability of graph query, reduces the data loading overhead and preprocessing overhead at the same time, and significantly improves the performance of the present invention.
[0074] The present invention conducts relevant experiments on the execution time. In the experimental configuration, an Alibaba Cloud ECS cluster is used: there are a total of six machines with the same configuration, ecs.g5.2xlarge instances (8 VCPUs, 32 GB of memory). Among them, two machines are storage nodes and management nodes, and the remaining four are computing nodes. Initially, only one computing node is used, and subsequent computing nodes are added to the cluster through elastic expansion according to the task requirements.
[0075] First, introduce the compared graph query systems: Gemini, Grape, and D-Galois are all current mainstream distributed graph query inventions. Gemini has extremely high computing efficiency by using a graph partitioning strategy with low overhead and an efficient graph compression method; Grape can automatically convert sequentially executed graph algorithms into parallel execution while retaining the optimization methods of sequential algorithms, and can simultaneously leverage the performance advantages of parallel and optimization methods; D-Galois supports different programming models and graph partitioning strategies, reduces the communication overhead of this invention by using structural and temporal invariance, and adopts a computing method of asynchronous computing inside computing nodes and synchronous computing between nodes, achieving a balance between computing throughput and communication overhead.
[0076] This invention changes the compared graph query system into a memory-computation separation architecture to evaluate the performance of traditional graph query systems under the memory-computation separation architecture. This invention compares the end-to-end running time of the existing mainstream graph query systems when executing graph query tasks, that is, the time from the start to the completion of the computing task. The end-to-end time of the traditional graph query system includes data loading time, preprocessing time, and graph query time. Figure 3 Shows the end-to-end time of this invention and Gemini, Grape, and D-Galois when executing the SSSP algorithm on different graphs. On most datasets, the performance of this invention far exceeds that of the existing systems, with a performance improvement of more than 10 times. Since the HW dataset is approximately a fully connected graph and cannot leverage the advantage of on-demand loading, we still have comparable performance with the existing systems.
[0077] The above description is only the preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. An elastic expansion graph query method with on-demand loading, characterized in that, including Step 1: For a directed acyclic graph, construct an improved CSR index file. The improved CSR index file includes an index file and an edge file. The edge file contains multiple lines of data. Each line of data includes the vertex pointed to in the directed acyclic graph, or each line of data includes the vertex pointed to in the directed acyclic graph and the weight of the edge between the pointing vertex and the pointed vertex. The index file includes the pointing vertex and the starting line number of the pointing vertex in the edge file; Step 2: Input the storage path of the index file and the storage path of the edge file at a pre-configured storage node. The storage node reads the index file from the local disk according to the storage path of the index file, constructs an index array for each vertex based on the index file. The index array of the vertex includes the starting line number of the vertex in the edge file. The storage node sends the number of vertices to the computing node and sends the storage location of the vertex to the management node; Step 3: A pre-configured management node receives the locally active graph algorithm to be executed input by the user and the source point in the directed acyclic graph. The locally active graph algorithm is an algorithm that can complete the execution task only with partial data of the directed acyclic graph; Take the source point as the current active vertex. The management node sends the data loading task of the current active vertex to the storage node and sends the current active vertex and the locally active graph algorithm to a pre-configured current computing node; Step 4: After receiving the data loading task of the current active vertex, the storage node reads the edge file from the local disk according to the storage path of the edge file; According to the index array of the current active vertex, obtain the edge data of multiple pointed vertices corresponding to the current active vertex in the edge file, that is, the edge data of multiple target vertices, send the edge data to the current computing node, and pre-read the edge data of the pointed vertices corresponding to the target vertices; Step 5: The current computing node receives all the current active vertices and the locally active graph algorithm, and initializes the state values of all vertices in the directed acyclic graph according to the algorithm type of the locally active graph algorithm; Receive the edge data corresponding to the current active vertex sent by the storage node, update the state values of all target vertices according to the locally active graph algorithm, take the target vertices as the new current active vertices, and obtain multiple current active vertices; Step 6: Determine the size of the edge data corresponding to all current active vertices, and obtain the memory capacity of the current computing node. According to the size of the edge data corresponding to all current active vertices and the obtained memory capacity of the current computing node, determine the target computing node; Step 7: The target computing node sends all current active vertices to the management node, and at the same time determines whether the edge data of the target vertex corresponding to the current active vertex has been stored in the target computing node among all current active vertices. If the edge data of the target vertex corresponding to the current active vertex has been stored in the target computing node among all current active vertices, the status value of the target vertex corresponding to the current active vertex is updated according to the local active graph algorithm. If the edge data of the target vertex corresponding to the current active vertex has not been stored in the target computing node among all current active vertices, the target computing node remains in a waiting state; Step 8: The management node receives all current active vertices. The management node sends the data loading task of the current active vertices to the storage node. After receiving the data loading task of the current active vertices, the storage node obtains the edge data of multiple pointed-to vertices of the current active vertices read in advance, that is, the edge data of multiple target vertices, and sends the edge data of multiple target vertices to the target computing node. Then, by the same token, the edge data of the pointed-to vertices of the target vertex is read in advance; Step 9: After receiving the edge data of multiple target vertices, the target computing node determines that there is no target vertex that has just been updated in Step 7 among all target vertices, and updates its status value according to the local active graph algorithm; Step 10: Determine whether the target vertex meets the termination condition. If the target vertex meets the termination condition, store the final status values of all vertices from the source vertex to the target vertex. If the target vertex does not meet the termination condition, return all updated target vertices as current active vertices and execute Step 6; 2. The elastic expansion graph query method with on-demand loading according to claim 1, wherein In Steps 3 and 8, the management node sends the data loading task of the current active vertices to the storage node, including: The management node determines the storage node storing the current active vertices and sends the current active vertices to the corresponding storage node; 3. The on-demand loading elastic expansion graph query method according to claim 1, characterized in that, In Step 4, obtaining the edge data of multiple pointed-to vertices corresponding to the current active vertex in the edge file includes: Obtain the index array of the current active vertex, and according to the starting line number of the current active vertex in the edge file in the index array, obtain the edge data of multiple pointed-to vertices corresponding to the current active vertex in the edge file; 4. A method for querying an on-demand loading elastic expansion graph according to claim 1, characterized in that, In Step 4, when the edge file contains the pointed-to vertices in the directed acyclic graph, the edge data of the target vertex only includes the target vertex pointed to by the current active vertex. When the edge file contains the pointed-to vertices in the directed acyclic graph, and the weights of the edges between the pointing vertex and the pointed-to vertex, the edge data of the target vertex includes the target vertex pointed to by the current active vertex, and the weight of the edge from the current active vertex to the target vertex; 5. A method for querying an on-demand loading elastic expansion graph according to claim 1, characterized in that The local active graph algorithm includes the shortest path algorithm SSSP, the breadth-first search algorithm BFS, the personalized page rank algorithm PPR, the penalty hit probability algorithm PHP, and the single-source widest path algorithm SSWP; 6. The on-demand loading elastic expansion graph query method according to claim 1, characterized in that Step 6 specifically includes: Determine the size of the edge data corresponding to all currently active vertices, and obtain the memory capacity of the current computing node. Determine whether the memory capacity of the current computing node is less than the size of the edge data corresponding to all currently active vertices. In the case where the memory capacity of the current computing node is less than the size of the edge data corresponding to all currently active vertices, the current computing node sends an expansion signal to the management node. The management node receives the expansion signal, expands new computing nodes, and the management node sends the routing information of the already loaded vertices to the new computing nodes, that is, the computing nodes where the already loaded vertices are located. The current computing node sends the final state values of all vertices from the source point to the currently active vertices stored locally to the new computing nodes, uses the new computing nodes as the target computing nodes, and executes step 7. In the case where the memory capacity of the current computing node is greater than or equal to the size of the edge data corresponding to all currently active vertices, use the current computing node as the target computing node and execute step 7.
7. A method for querying an on-demand loading elastic expansion graph according to claim 1, characterized in that, The termination conditions described in step 10 include that the amount of state change is less than the termination threshold or the state no longer changes.