A Distributed-Oriented Graph Computing Implementation Method and System

By adopting a data partitioning method combining edge slicing and point slicing in a distributed system, combined with GAS computing model and buffer optimization technology, the problem of slow data processing speed of super-large-scale graph structures is solved, and more efficient computing efficiency is achieved.

CN117591510BActive Publication Date: 2025-06-20XIAMEN YUANTING INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311423995.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-30
Publication Date
2025-06-20
Estimated Expiration
2043-10-30

AI Technical Summary

Technical Problem

The prior art has slow calculation speed when processing ultra-large-scale graph structure data, especially when the graph data has Power Law properties, the graph division method leads to unbalanced load of computing nodes, affecting computing efficiency.

Method used

A distributed graph calculation implementation method is adopted, and data partitioning is processed through a combination of edge slicing and point slicing, calculation is performed using the GAS calculation model, and data transmission and storage are optimized through buffer and data compression.

Benefits of technology

The processing efficiency of large-scale graph structure data is improved and the problem of slow computing speed is solved. Especially when graph data has Power Law properties, it significantly improves computing efficiency by balancing load and optimizing data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117591510B_ABST
    Figure CN117591510B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for implementing graph computing for distribution. The method includes: S1: obtaining graph-format data corresponding to data to be analyzed, and extracting each vertex and relationship data between each vertex included therein; S2: performing data partitioning processing on the graph-format data based on each vertex and the relationship data between each vertex; S3: mapping all partitioning results corresponding to the graph-format data to each computing node in the distributed system, and sending the partitioning results and mapping relationships to each computing node; S4: calculating each received partitioning result by a GAS computing model at the computing node. The present invention improves the processing efficiency of large-scale graph-structured data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of graph computing, and particularly to a method and system for implementing distributed graph computing. Background Art

[0002] In recent years, with the advent of the Internet era, the amount of data humans possess has become increasingly huge. Ordinary single-machine graph computing frameworks are increasingly difficult to process this data. The adoption of distributed computing frameworks has increasingly become the primary choice for many enterprises, such as MapReduce, GraphX, etc. Some of these computing frameworks are not specifically developed for graph data formats, so some difficult problems in graph domain computing have not been well solved, and sometimes their performance in terms of computing speed is not good in certain special scenarios.

[0003] In recent years, many excellent graph computing frameworks / calculation models have been developed by many researchers. For example, Giraph is a typical distributed computing system, which is implemented with reference to Google's Pregel paper and is also called a programming model. This programming model adopts the BSP computing mode (Bulk Synchronous Parallel) and advocates the idea of "Think Like a Vertex". This model can be described as "each calculation consists of a series of supersteps, and each superstep can be divided into three steps: local concurrent calculation, global communication, and synchronization". The advantage of this mode is that there is no communication consumption during local calculation and no concurrent control is required. However, the disadvantage of BSP is that the overhead during synchronization is large, and if the load is not balanced, it is easy to cause a slow execution of a certain computing node, dragging down the execution efficiency of the entire superstep. The Pregel system adopts a vertex-centered partitioning method for graph partitioning. This method can well partition vertices to different computing nodes, and the edges of the vertices will also be partitioned to the same computing node along with the vertices. If the graph data has the Power Law property, this kind of computing partitioning method does not perform well during calculation. It can be imagined that if a vertex has millions or even tens of millions of edges and is all allocated to the same computing node, the computing volume and communication volume of this computing node will be huge, and all other nodes in the cluster have to wait for this computing node to complete, which will undoubtedly affect the calculation time. Subsequently, researchers have also proposed other calculation models and computing modes to address the above problems, such as GAS. This computing mode adopts a two-dimensional graph partitioning method. We generally refer to the graph partitioning method based on vertex partitioning as one-dimensional partitioning, and the two-dimensional partitioning method adopts an edge / relationship-based partitioning method. This method can evenly allocate edges to different computing nodes, thereby achieving the purpose of computing balance. However, this method will generate a large number of vertex copy data when partitioning edges to different computing nodes, and these copies will generate communication overhead and also have a certain impact on the calculation time.

[0004] In addition to the difficult problem of graph partitioning, some characteristics of graph data, such as random access, have a crucial impact on improving the efficiency of graph computing by enhancing the utilization rate of the cache. Also, due to the storage of massive data, it is impossible to store all the data in memory. If data needs to be read from disk or other storage media, it is necessary to improve the data reading speed and avoid random access to the storage media. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a distributed-oriented graph computing implementation method and system.

[0006] The specific solutions are as follows:

[0007] A method for implementing distributed graph computing, comprising the following steps:

[0008] S1: Obtain graph format data corresponding to the data to be analyzed, and extract each vertex and the relationship data between each vertex included therein;

[0009] S2: Based on each vertex and the relationship data between each vertex, perform data partitioning processing on the graph format data; in data partitioning, first perform a primary partitioning in the way of edge splitting, and then judge whether the number of edges included in each primary partitioning result is greater than a preset number threshold. If so, perform a secondary partitioning on the primary partitioning result in the way of vertex splitting;

[0010] S3: Map all partitioning results corresponding to the graph format data to each computing node in the distributed system, and send the partitioning results and the mapping relationship to each computing node;

[0011] S4: The computing node calculates each received partitioning result through the GAS computing model. When calculating each partitioning result to the Gatter stage, judge whether the partitioning result is a partitioning result that has undergone secondary partitioning. If so, when the Gatter stage calculation is completed, judge whether the Gatter stage of other secondary partitioning results mapped to the computing nodes belonging to the same primary partitioning result as this partitioning result has been calculated. If the calculation is completed, pull the Gatter stage calculation results of other secondary partitioning results and merge them with the Gatter stage calculation results of this partitioning result, and synchronize the merged result to the computing nodes mapped by other secondary partitioning results, so that the vertex data is updated based on the merged result in the Apply stage; if the calculation is not completed, wait for the merged result synchronized by other computing nodes.

[0012] Further, the method for sending the partitioning result to each computing node in step S3 is: construct a buffer for each computing node, and write the partitioning result into the buffer of each computing node according to the mapping relationship; judge in real time whether the amount of data stored in each buffer reaches the buffer size limit. If so, notify the computing node corresponding to the buffer. After receiving the notification, the computing node pulls the data in the buffer into the computing node.

[0013] Further, when writing the partitioning result into the buffer, write it into each buffer in sequence according to the vertex number.

[0014] Further, perform a compression operation on the data to be pulled before the data in the buffer is pulled.

[0015] Further, after the computing node pulls the data, store the pulled data in the memory or block disk of the computing node.

[0016] A distributed graph computing implementation system includes terminal devices and computing nodes under a distributed system, and the system implements the steps of the method in the above embodiments of the present invention.

[0017] By adopting the above technical solution, the present invention solves the problem of slow data processing speed for ultra-large-scale graph-structured data in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 The flowchart of Embodiment 1 of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To further illustrate the embodiments, the present invention provides drawings. These drawings are part of the disclosure of the present invention, mainly used to illustrate the embodiments, and can be combined with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention.

[0020] The present invention will be further described below in conjunction with the drawings and specific embodiments.

[0021] Embodiment 1:

[0022] The embodiment of the present invention provides a distributed graph computing implementation method. The graph data types that can be used are: financial field data, applied to technical problems such as risk control, anti-fraud (mining customer credit risk levels), business association analysis under large data volumes, and equity penetration; public security field data, applied to personnel relationship network analysis (constructing a large relationship network of people, events, things, organizations, and virtual identities); social network data, applied to technical problems such as network density analysis, network path betweenness analysis, and community discovery analysis.

[0023] As Figure 1 shown, the method includes the following steps:

[0024] S1: Obtain graph format data corresponding to the data to be analyzed, and extract each vertex and the relationship data between each vertex included therein.

[0025] S2: Based on each vertex and the relationship data between each vertex, perform data partitioning processing on the graph format data. In data partitioning, first perform a primary partitioning in the way of edge splitting (based on a one-dimensional graph partitioning method), and then determine whether the number of edges included in each primary partitioning result is greater than a preset number threshold. If so, perform a secondary partitioning on the primary partitioning result in the way of vertex splitting (based on a two-dimensional graph partitioning method).

[0026] Those skilled in the art can set the size of the quantity threshold by themselves, which is set to 200 in this embodiment. After the first partition, different vertices are in different partition results. After the second partition, the same vertex may be in different partition results.

[0027] S3: Map all partition results corresponding to the graph format data to each computing node in the distributed system, and send the partition results and the mapping relationship to each computing node.

[0028] The partition result during mapping may be the first partition result with the number of included edges not greater than the preset quantity threshold, or may be the second partition result obtained by further partitioning the first partition result with the number of included edges not greater than the preset quantity threshold.

[0029] The mapping relationship includes which vertices are included in the partition results mapped by each computing node.

[0030] Further, in this embodiment, when sending the partition results to each computing node, a buffer is constructed for each computing node, and the partition results are written into the buffers of each computing node according to the mapping relationship; it is judged in real time whether the amount of data stored in each buffer reaches the buffer size limit. If so, the computing node corresponding to the buffer is notified. After receiving the notification, the computing node pulls the data in the buffer into the computing node.

[0031] The size of the buffer can be set according to the memory resources of the computing node.

[0032] In this embodiment, when writing the partition results into the buffer, they can be written into each buffer in the order of vertex numbers. The purpose of sorting is to write to the disk sequentially. Compared with the random reading method, it can speed up the I / O reading and writing speed.

[0033] Before the data in the buffer is pulled, a compression operation needs to be performed on the data to be pulled to reduce the communication time consumption and improve the computing efficiency.

[0034] After the computing node pulls the data, the pulled data is stored in the memory or the partitioned disk of the computing node for calculation. The purpose of partitioning is to speed up the I / O reading and writing speed and improve the computing efficiency. The pulling method is used instead of the data pushing method to reduce the blocking of the local partitioning thread and improve the partitioning efficiency.

[0035] S4: The computing node calculates the received partition results through the GAS computing model.

[0036] Each time the GAS computing model processes graph data, it is divided into three stages: the Gatter (information collection) stage, the Apply (updating vertex information) stage, and the Scatter (updating adjacent edges and vertex information) stage. The Gatter stage is responsible for collecting information on adjacent edges and vertices, and then running a user-defined function for aggregation calculation. The Apply stage is responsible for updating the results of the previous aggregation calculation to the corresponding vertices. The Scatter stage updates the vertex information in the Apply stage to the adjacent edges and vertices.

[0037] When calculating each partition result to the Gatter stage, it is determined whether the partition result is a result of secondary partitioning. If so, when the Gatter stage calculation is completed, it is determined whether the Gatter stage of other secondary partition results mapped to the computing nodes of the same primary partition result as this partition result has been calculated. If the calculation is completed, the Gatter stage calculation results of other secondary partition results are pulled and merged with the Gatter stage calculation results of this partition result, and the merged result is synchronized to the computing nodes mapped by other secondary partition results, so that the vertex data is updated based on the merged result in the Apply stage; if the calculation is not completed, wait for the merged result synchronized by other computing nodes. If the partition result is a primary partition result, it directly enters the Apply stage when the Gatter stage calculation is completed.

[0038] The embodiments of the present invention improve the processing efficiency of large-scale graph structure data.

[0039] Embodiment 2:

[0040] The present invention also provides a graph computing implementation system for distributed systems, including a terminal device and computing nodes under a distributed system. The system implements the steps in the above method embodiments of Embodiment 1 of the present invention.

[0041] The terminal device is used to perform the partitioning and mapping of graph format data in steps S1 to S3, and then implement the calculation operation in step S4 through the computing nodes.

[0042] Although the present invention has been specifically shown and described in conjunction with the preferred embodiments, those skilled in the art should understand that various changes can be made to the present invention in form and detail without departing from the spirit and scope of the present invention defined by the appended claims, and all of them fall within the protection scope of the present invention.

Claims

1. A method for implementing distributed graph computing, characterized in that: It includes the following steps: S1: Obtain the graph format data corresponding to the data to be analyzed, and extract the relationship data between each vertex and each vertex contained therein; S2: Based on the relationship data between each vertex and each vertex, perform data partitioning processing on the graph format data; in data partitioning, first perform a primary partitioning by means of edge splitting, and then determine whether the number of edges contained in each primary partitioning result is greater than a preset number threshold. If so, perform a secondary partitioning on the primary partitioning result by means of vertex splitting; S3: Map all partitioning results corresponding to the graph format data to each computing node in the distributed system, and send the partitioning results and the mapping relationship to each computing node; S4: The computing node calculates each received partitioning result through the GAS computing model. When calculating each partitioning result to the Gatter stage, determine whether the partitioning result is a partitioning result that has undergone secondary partitioning. If so, when the Gatter stage calculation is completed, determine whether the Gatter stage of other secondary partitioning results mapped to the computing nodes belonging to the same primary partitioning result as this partitioning result has been calculated. If the calculation is completed, pull the Gatter stage calculation results of other secondary partitioning results and merge them with the Gatter stage calculation results of this partitioning result, and synchronize the merged result to the computing nodes mapped by other secondary partitioning results, so that the vertex data is updated based on the merged result in the Apply stage; If the calculation is not completed, wait for the merged result synchronized by other computing nodes.

2. The method for implementing distributed graph computing according to claim 1, characterized in that: The method of sending the partitioning results to each computing node in step S3 is: construct a buffer for each computing node, and write the partitioning results into the buffers of each computing node according to the mapping relationship; continuously judge whether the amount of data stored in each buffer reaches the buffer size limit. If so, notify the computing node corresponding to the buffer. After receiving the notification, the computing node pulls the data in the buffer into the computing node.

3. The method for implementing distributed graph computing according to claim 2, characterized in that: When writing the partitioning results into the buffer, write them into each buffer in the order of the vertex numbers.

4. The method for implementing distributed graph computing according to claim 2, characterized in that: Perform a compression operation on the data to be pulled before the data in the buffer is pulled.

5. The method for implementing distributed graph computing according to claim 2, characterized in that: After the computing node pulls the data, store the pulled data in the memory or chunk disk of the computing node.

6. A system for implementing distributed graph computing, characterized in that: It includes a terminal device and a computing node under a distributed system. The system implements the method described in any one of claims 1 to 5. Specifically, after the terminal device is used for the partitioning and mapping work of the graph format data in steps S1 to S3, the computing operation in step S4 is implemented through the computing node.

Citation Information

Patent Citations

  • Parallel constraint subgraph mining method based on edge-node mixed segmentation

    CN114722241A