Graph data division method and device, graph data processing method and device, equipment and medium
By determining the node with the largest degree in graph data processing for division and adjusting the sub-graph according to the number of edge nodes, the problem of unreasonable sub-graph division is solved, and more efficient graph processing performance is achieved.
Patent Information
- Application Number
- CN202510137663.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-16
AI Technical Summary
When using a computing model centered on subgraphs for single-machine large-scale graph analysis, unreasonable subgraph division leads to low memory usage efficiency, large disk I/O overhead and low parallelism of computing tasks, affecting the overall graph processing performance.
By determining the node with the largest degree from the target graph data as the target node, the target graph data is divided to obtain the initial subgraph, and adjust it according to the number of edge nodes in the initial subgraph to obtain the updated subgraph. The storage capacity of the updated subgraph must be in the preset range, otherwise the division process will be repeated until all subgraphs are suitable for storage.
The classification strategy of graph data is optimized, the performance of graph processing is improved, and by maintaining connectivity and load balancing between sub-graphs, communication overhead and storage optimization are reduced, and the convergence efficiency of graph processing is improved.
Smart Images

Figure CN120011603A_ABST
Abstract
Description
Technical Field
[0001] The present application is applicable to the field of data processing, and in particular relates to a graph data partitioning method, a graph data processing method, a device, a equipment and a medium. Background Art
[0002] In the era of big data, graph data, as the carrier of complex information networks, is becoming increasingly large in scale, and the demand for graph data processing and analysis has also increased sharply. Traditional large-scale graph computing systems use a parallel method of data partitioning, that is, integrating multiple computer resources to complete graph computing tasks. Considering the high maintenance and construction costs, a series of large-scale graph processing systems based on single machines have emerged. These systems use disks as an extension of memory to process large graphs, and adopt a vertex-centric computing model (this model limits the information in the computing process to be transmitted between nodes) or a subgraph-centric computing model (this model transmits the information in the computing process between subgraphs) for calculation. However, although large-scale graph processing systems based on single machines have made certain progress in processing large graph data, when using a subgraph-centric computing model and using disks as external storage for graph data to implement large-scale graph analysis tasks on a single machine, there is still an unreasonable subgraph partitioning situation. This irrationality affects the efficiency of memory use, increases the overhead of disk I / O, and reduces the parallelism of computing tasks, thereby affecting the overall performance of graph processing. Therefore, how to optimize the partitioning strategy of graph data to improve the performance of graph processing has become an urgent problem to be solved. Summary of the invention
[0003] In view of this, embodiments of the present application provide a graph data partitioning method, a graph data processing method, an apparatus, a device and a medium to solve the problem of how to optimize the graph data partitioning strategy to improve the performance of graph processing.
[0004] In a first aspect, an embodiment of the present application provides a graph data partitioning method, the graph data partitioning method comprising: Obtain target graph data, determine the out-degree of each node in the target graph data, and determine the node with the largest out-degree from the target graph data as the target node; According to the target node, the target graph data is divided to obtain at least two initial subgraphs, and edge nodes in each initial subgraph are determined, wherein each initial subgraph includes the target node, and the edge node refers to a node included in at least two initial subgraphs; Obtaining the number of edge nodes in all initial subgraphs, adjusting the node paths corresponding to the edge nodes in at least one initial subgraph with the goal of ensuring that the number of edge nodes is within a preset value range, and obtaining an updated subgraph; For any updated subgraph, if the storage capacity required to store the updated subgraph is not within the preset capacity range, the updated subgraph is replaced by the target graph data, and the step of determining the out-degree of each node in the target graph data is returned to execute until all updated subgraphs are obtained, wherein the updated subgraph is used for partitioned storage on the disk.
[0005] In a second aspect, an embodiment of the present application provides a graph data processing method. After all updated subgraphs are obtained by using the graph data partitioning method of the first aspect, the graph data processing method includes: In one round of iterative processing, all updated subgraphs in the partition are read from the partition of the disk, and all updated subgraphs in the partition are transferred to the memory buffer; Monitoring the memory capacity occupied by the updated subgraph in the memory buffer, and when the memory capacity reaches a preset threshold, transmitting the updated subgraph in the memory buffer to the memory execution area; According to a preset graph algorithm, the update subgraphs in the memory execution area are processed in parallel, and according to the processing process, the state of each update subgraph is updated; For any updated subgraph, if the state of the updated subgraph reaches a preset termination state, the processing of the updated subgraph is stopped to obtain a processed updated subgraph, and the processed updated subgraph is written to the disk.
[0006] In a third aspect, an embodiment of the present application provides a graph data partitioning device, the graph data partitioning device comprising: An acquisition module is used to acquire target graph data, determine the out-degree of each node in the target graph data, and determine the node with the largest out-degree as the target node from the target graph data; A partitioning module is used to partition the target graph data according to the target node to obtain at least two initial subgraphs, and determine the edge nodes in each initial subgraph, wherein each initial subgraph includes the target node, and the edge nodes refer to nodes included in at least two initial subgraphs; An adjustment module is used to obtain the node number of edge nodes in all initial subgraphs, and to adjust the node path corresponding to the edge node in at least one initial subgraph with the node number of the edge node being within a preset value range as a goal, so as to obtain an updated subgraph; The first loop module is used to, for any updated subgraph, if the storage capacity required to store the updated subgraph is not within a preset capacity range, then the updated subgraph is used as the target graph data, and the step of determining the out-degree of each node in the target graph data is returned to execute until all the updated subgraphs are obtained, wherein the updated subgraphs are used for partitioned storage in the disk.
[0007] In a fourth aspect, an embodiment of the present application provides a graph data processing device. After all updated subgraphs are obtained by using the graph data partitioning method of the first aspect, the graph data processing device includes: A reading module, used for reading all updated subgraphs in the partition from the partition of the disk in a round of iterative processing, and transferring all updated subgraphs in the partition read to a memory buffer; A scheduling module, used to monitor the memory capacity occupied by the updated subgraph in the memory buffer, and when the memory capacity reaches a preset threshold, transfer the updated subgraph in the memory buffer to the memory execution area; An operation module, configured to perform parallel processing on the update subgraphs in the memory execution area according to a preset graph algorithm, and update the state of each update subgraph according to the processing process; The writing module is used for stopping processing any updated subgraph if the state of the updated subgraph reaches a preset termination state, obtaining a processed updated subgraph, and writing the processed updated subgraph into the disk.
[0008] In a fifth aspect, an embodiment of the present application provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the graph data partitioning method as described in the first aspect or the graph data processing method as described in the second aspect is implemented.
[0009] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the graph data partitioning method as described in the first aspect, or the graph data processing method as described in the second aspect is implemented.
[0010] Compared with the prior art, the embodiments of the present application have the following beneficial effects: the graph data partitioning method of the present application determines the node with the largest out-degree as the target node from the target graph data, partitions the target graph data according to the target node, obtains at least two initial subgraphs, determines the edge nodes in each initial subgraph, each initial subgraph includes the target node, and the edge node refers to the node included in at least two initial subgraphs, obtains the number of edge nodes in all initial subgraphs, and adjusts the node paths corresponding to the edge nodes in at least one initial subgraph with the goal of the number of edge nodes being within a preset numerical range to obtain an updated subgraph. For any updated subgraph, if the storage capacity required to store the updated subgraph is not within the preset capacity range, the updated subgraph is used as the target graph data, and returns to execute the step of determining the out-degree of each node in the target graph data until all updated subgraphs are obtained.
[0011] The graph data processing method of the present application transfers the updated subgraph read from the partition of the disk to the memory buffer in a round of iterative processing. When the content capacity occupied by the updated subgraph in the memory buffer reaches a preset threshold, the updated subgraph in the memory buffer is transferred to the memory execution area. According to the preset graph algorithm, the updated subgraphs in the memory execution area are processed in parallel, and the state of each updated subgraph is updated according to the processing process. For any updated subgraph, if the state of the updated subgraph reaches a preset termination state, the processing of the updated subgraph is stopped to obtain the processed updated subgraph, and the processed updated subgraph is written to the disk.
[0012] Among them, by selecting the node with the largest out-degree as the target node, and taking this as the benchmark to divide the target graph data, the initial subgraph is obtained, which ensures the connectivity between the initial subgraphs and avoids separating highly connected nodes, thereby reducing the communication overhead in subsequent graph processing; on this basis, based on the number of edge nodes in all initial subgraphs, the initial subgraph is adjusted to obtain an updated subgraph, and on the basis of connectivity, the load balancing of nodes or edges between each updated subgraph is achieved; and based on the storage capacity required by the subgraph, the updated subgraph is adjusted to optimize the storage and processing performance of the subgraph, thereby optimizing the overall division strategy of the graph data; therefore, when the updated subgraphs obtained by the above division are processed in parallel based on the preset graph algorithm, since the updated subgraphs maintain good connectivity, the information synchronization speed between the updated subgraphs is greatly promoted. During the processing, information can be transmitted more quickly between the updated subgraphs, which improves the convergence efficiency of the graph processing and thus improves the performance of the graph processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0014] Figure 1 This is a schematic diagram of an application environment of a graph data partitioning method and a graph data processing method provided in Example 1 of the present application; Figure 2 It is a flowchart of a graph data partitioning method provided in Embodiment 2 of the present application; Figure 3 is a schematic diagram of target graph data provided in Example 2 of the present application; Figure 4 is a schematic diagram of updating a subgraph provided in Embodiment 2 of the present application; Figure 5It is a flowchart of a graph data partitioning method provided in Embodiment 3 of the present application; Figure 6 It is a flowchart of a graph data processing method provided in Embodiment 4 of the present application; Figure 7 This is a schematic diagram of a system architecture for graph data processing in a single-machine multi-core environment provided in Embodiment 4 of the present application; Figure 8 It is a flowchart of a graph data processing method provided in Embodiment 5 of the present application; Fig. 9 It is a structural schematic diagram of a graph data partitioning device provided in Embodiment 6 of the present application; Fig.10 It is a structural schematic diagram of a graph data processing device provided in Embodiment 7 of the present application; Fig.11 It is a structural diagram of a computer device provided in Example 8 of the present application. DETAILED DESCRIPTION
[0015] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0016] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0017] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0018] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0019] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0020] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0021] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0022] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0023] It should be understood that the size of the serial numbers of the steps in the following embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0024] In order to illustrate the technical solution of the present application, a specific embodiment is provided below for illustration.
[0025] A graph data partitioning method and a graph data processing method provided in the first embodiment of the present application can be applied in the following aspects: Figure 1In the application environment, the server communicates with the client, the server provides graph data partitioning service and graph data processing service, and the client triggers graph data partitioning task and graph data processing task to the server. The client includes but is not limited to PDA, desktop computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, cloud computer equipment, personal digital assistant (PDA) and other devices. The computer equipment corresponding to the server can be implemented by an independent server or a server cluster composed of multiple servers.
[0026] See also Figure 2 , is a flow chart of a graph data partitioning method provided in Embodiment 2 of the present application, and the graph data partitioning method is applied to Figure 1 The server in the example connects to the client to obtain the target graph data sent by the client. Figure 2 As shown, the graph data partitioning method may include the following steps: Step S201, obtain target graph data, determine the out-degree of each node in the target graph data, and determine the node with the largest out-degree from the target graph data as the target node.
[0027] Step S202: divide the target graph data according to the target node to obtain at least two initial subgraphs, and determine the edge nodes in each initial subgraph.
[0028] In this embodiment, the target graph data may refer to the data of the graph structure to be divided, the target node may refer to the node with the largest out-degree in the target graph data, the initial subgraph may refer to the graph data obtained by dividing the target graph data according to the target node, each initial subgraph includes the target node, and the edge node refers to the node included in at least two initial subgraphs.
[0029] Specifically, for any node in the target graph data, determine the number of edges emitted from the node, select the node with the largest number of edges emitted from the target graph data as the target node, divide the target graph data starting from the target node to obtain at least two initial subgraphs, each of which includes the target node, and for any initial subgraph, determine the edge nodes in the initial subgraph that are included in at least two initial subgraphs. Since each initial subgraph includes the target node, the target node is an edge node in each initial subgraph.
[0030] Step S203, obtaining the node numbers of edge nodes in all initial subgraphs, adjusting the node paths corresponding to the edge nodes in at least one initial subgraph with the node number of the edge nodes being within a preset value range as a goal, and obtaining an updated subgraph.
[0031] In this embodiment, the node path may refer to a path containing edge nodes, and the updated subgraph may refer to graph data in which the number of edge nodes obtained after adjusting the node paths corresponding to the edge nodes in at least one initial subgraph is within a preset value range.
[0032] Specifically, the number of edge nodes in all initial subgraphs is obtained. If the number of edge nodes in all initial subgraphs is already within a preset value range, each initial subgraph can be determined to be an updated subgraph. If the node number of edge nodes in all initial subgraphs is not within a preset numerical range, then for any initial subgraph, the edge nodes in the initial subgraph are determined, and for any edge node in the initial subgraph except the target node, the node path containing the edge node is determined in the initial subgraph, and for any node path, the node path is randomly added to any initial subgraph including the edge node except the initial subgraph, and the edge nodes are merged to reduce the node number of all edge nodes. When the node number of all edge nodes is within the preset numerical range, an updated subgraph is obtained.
[0033] Step S204, for any updated subgraph, if the storage capacity required to store the updated subgraph is not within the preset capacity range, the updated subgraph is set as the target graph data, and the step of determining the out-degree of each node in the target graph data is returned to execute until all updated subgraphs are obtained.
[0034] In this embodiment, the storage capacity may refer to the storage space required for updating the sub-graph in the storage device, and may be measured using basic units of data storage capacity such as kilobytes (KB) and megabytes (MB).
[0035] For example, if the preset capacity range is at the KB-MB level, for any updated sub-graph, if the storage capacity required to store the updated sub-graph is at the KB-MB level, there is no need to further divide the updated sub-graph; if the storage capacity required to store the updated sub-graph is not at the KB-MB level, it is necessary to further divide the updated sub-graph, and use the updated sub-graph as the new target graph data. Return to execute the steps S201 to S203 above until the storage capacity required for all updated sub-graphs is at the KB-MB level, and all updated sub-graphs are obtained, and the obtained updated sub-graph partitions are stored on the disk.
[0036] For example, see Figure 3 , is a schematic diagram of a target graph data provided in Example 2 of the present application. Figure 3As shown, there are 10 nodes in the target graph data, namely node 0, node 1, node 2, node 3, node 4, node 5, node 6, node 7, node 8 and node 9. The out-degree of node 0 is 5, the out-degree of node 1 is 2, the out-degree of node 2 is 2, the out-degree of node 3 is 3, the out-degree of node 4 is 3, the out-degree of node 5 is 3, the out-degree of node 6 is 3, the out-degree of node 7 is 1, the out-degree of node 8 is 1, and the out-degree of node 9 is 3. It can be determined that the node with the largest out-degree in the target graph data is node 0, that is, node 0 is determined as the target node.
[0037] See also Figure 4 , is a schematic diagram of an updated subgraph provided in Embodiment 2 of the present application, wherein the updated subgraph is Figure 3 The updated subgraph obtained by dividing the target graph data shown in Figure 4 As shown, Figure 3 The target graph data shown is divided into 5 update subgraphs, namely update subgraph B1, update subgraph B2, update subgraph B3, update subgraph B4 and update subgraph B5. The update subgraphs are interconnected, the number of edge nodes in all update subgraphs is within a preset value range, the difference in the total number of corresponding nodes between all update subgraphs is less than a preset value, and the storage capacity required to store each update subgraph is also within a preset capacity range.
[0038] In the embodiment of the present application, the node with the largest out-degree is selected as the target node, and the target graph data is divided based on this to obtain the initial subgraph, thereby ensuring the connectivity between the initial subgraphs and avoiding separating highly connected nodes, thereby reducing the communication overhead in subsequent graph processing; on this basis, based on the number of edge nodes in all initial subgraphs, the initial subgraph is adjusted to obtain an updated subgraph, and on the basis of connectivity, load balancing of nodes or edges between each updated subgraph is achieved; and based on the storage capacity required by the subgraph, the updated subgraph is adjusted to optimize the storage and processing performance of the subgraph, thereby optimizing the overall division strategy of the graph data; therefore, when the updated subgraphs obtained by the above divisions are processed in parallel based on the preset graph algorithm, since the updated subgraphs maintain good connectivity, the information synchronization speed between the updated subgraphs is greatly promoted. During the processing, information can be transmitted more quickly between the updated subgraphs, which improves the convergence efficiency of the graph processing and thus improves the performance of the graph processing.
[0039] See also Figure 5 , is a flow chart of a graph data partitioning method provided in Example 3 of the present application. Figure 5As shown, in the above step S203, the node path corresponding to the edge node in at least one initial subgraph is adjusted with the goal of the node number of the edge node being within a preset value range to obtain an updated subgraph, which may include the following steps: Step S501, for any initial subgraph, determine the edge nodes in the initial subgraph, and for any edge node in the initial subgraph except the target node, determine the node path containing the edge node in the initial subgraph.
[0040] Step S502: for any node path, add the node path to any initial subgraph including edge nodes except the initial subgraph to obtain a subgraph to be evaluated.
[0041] Step S503, determining the total number of nodes in each subgraph to be evaluated and the number of edge nodes in all subgraphs to be evaluated, and for any subgraph to be evaluated, respectively calculating the difference between the total number of nodes in the subgraph to be evaluated and the total number of nodes in each subgraph to be evaluated except the subgraph to be evaluated.
[0042] Step S504: if the calculated differences corresponding to all subgraphs to be evaluated are smaller than a preset value, and the number of edge nodes in all subgraphs to be evaluated is within a preset value range, each subgraph to be evaluated is determined to be an updated subgraph.
[0043] In this embodiment, the subgraph to be evaluated may refer to graph data after the initial subgraph is updated according to the node path of the edge node.
[0044] Optionally, after respectively calculating the difference between the total number of nodes in the subgraph to be evaluated and the total number of nodes in each subgraph to be evaluated except the subgraph to be evaluated in the above step S503, the method further includes: If there is a subgraph to be evaluated whose corresponding difference is not less than a preset value, or the number of edge nodes in all subgraphs to be evaluated is not within the preset value range, then for any subgraph to be evaluated, the subgraph to be evaluated is set as the initial subgraph, and the step of determining the edge nodes in the initial subgraph is returned to execute until each subgraph to be evaluated is determined to be an updated subgraph.
[0045] That is, if there is a subgraph to be evaluated whose corresponding difference is not less than a preset value, or the number of edge nodes in all subgraphs to be evaluated is not within the preset value range, then for any subgraph to be evaluated, the subgraph to be evaluated is treated as a new initial subgraph, and the execution of the above steps S501 to S503 is returned until each subgraph to be evaluated is determined to be an updated subgraph.
[0046] For example, in the process of determining the updated subgraph, if the target graph data is divided, two initial subgraphs are obtained, namely initial subgraph A1 and initial subgraph A2, the number of edge nodes in the two initial subgraphs is 4, there are two edge nodes in the initial subgraph A1, namely node 0 and node 2, and there are two edge nodes in the initial subgraph A2, namely node 0 and node 2, where node 0 is the target node; In the initial subgraph A1, node 2 corresponds to a node path: node 2→node 4, and in the initial subgraph A2, node 2 corresponds to a node path: node 2→node 5; for the edge node: node 2 in the initial subgraph A2, if the node path corresponding to node 2: node 2→node 5 is added to another initial subgraph A1 including node 2, after adding, the subgraph A1 to be evaluated and the subgraph A2 to be evaluated are obtained. There are two paths starting from node 2 in the subgraph A1 to be evaluated, namely node 2→node 4 and node 2→node 5, and there is no path corresponding to node 2 and node 2 in the subgraph A2 to be evaluated. At this time, the number of edge nodes in the two subgraphs to be evaluated is 2; If the preset value range is [1,3], it can be determined that the number of edge nodes 2 in the two subgraphs to be evaluated is in the preset value range [1,3]; if the difference in the total number of nodes corresponding to the two subgraphs to be evaluated is also less than the preset value, it can be determined that the two subgraphs to be evaluated are both update subgraphs.
[0047] In an embodiment of the present application, the node paths corresponding to the edge nodes in at least one initial subgraph are adjusted with the goal of ensuring that the number of edge nodes is within a preset numerical range and that the differences corresponding to all subgraphs to be evaluated are less than a preset numerical value, thereby obtaining an updated subgraph. While ensuring the connectivity between the updated subgraphs, the load balance between the updated subgraphs is maintained, the storage and processing performance of the subgraphs is optimized, and when graph processing is performed based on the updated subgraphs, the convergence efficiency of the graph processing is improved, thereby improving the performance of the graph processing.
[0048] See also Figure 6 , is a flow chart of a graph data processing method provided in Example 4 of the present application. Figure 6 As shown, after all updated subgraphs are obtained by adopting the graph data partitioning method, the graph data processing method may include the following steps: Step S601, in a round of iterative processing, all updated subgraphs in a partition are read from the partition of the disk, and all updated subgraphs in the partition are transferred to the memory buffer.
[0049] Step S602 , monitoring the memory capacity occupied by the updated subgraph in the memory buffer, and when the memory capacity reaches a preset threshold, transmitting the updated subgraph in the memory buffer to the memory execution area.
[0050] Step S603 , according to a preset graph algorithm, the update subgraphs in the memory execution area are processed in parallel, and according to the processing process, the state of each update subgraph is updated.
[0051] Step S604: for any updated subgraph, if the state of the updated subgraph reaches a preset termination state, the processing of the updated subgraph is stopped to obtain a processed updated subgraph, and the processed updated subgraph is written to the disk.
[0052] In this embodiment, the memory buffer is used to temporarily store data read from an external device or a data source, the memory execution area is used to store the program instructions being executed and related data, and the preset graph algorithm may refer to a preset algorithm for processing the update subgraph, for example, the preset graph algorithm may be an algorithm for traversing, updating, and querying the update subgraph; The state of the update subgraph may refer to the processing state of the update subgraph. For example, the state of the update subgraph may include an activation state, a waiting state for calculation, a calculating state, and a converged state. The first two states indicate that the update subgraph is on disk, and the remaining states indicate that the update subgraph is in memory. The initial state of each update subgraph is an activation state, which means that the update subgraph is waiting to be read into memory; when the update subgraph is transferred to the memory buffer or the memory execution area, the state of the update subgraph is a waiting state for calculation; when the update subgraph is processed according to a preset graph algorithm, the state of the update subgraph is a calculating state; when the update subgraph converges during processing, the state of the update subgraph is a converged state; The preset termination state may refer to a preset state that the update subgraph needs to reach when stopping processing the update subgraph. For example, if it is necessary to detect whether the update subgraph has converged by the state of the update subgraph, the preset termination state may be set to the state corresponding to the state when the update subgraph converges. For example, when the state of the update subgraph is a converged state, the preset termination state has been reached.
[0053] This application follows the overall synchronization model for the processing of all updated subgraphs corresponding to the target graph data. In one round of iterative processing, all updated subgraphs in the partitions read from the disk partitions are calculated and converged, and then enter the next round of iterative processing after synchronization. That is, the calculations between the updated subgraphs are performed asynchronously, and the synchronization is controlled by the partition where the updated subgraph is located.
[0054] Specifically, a round of iterative processing can be divided into three periods: a read period, a run period, and a write period. The three periods work in a pipeline mode. In the read period, since the storage capacity of the updated subgraph obtained by dividing the target graph data using the above-mentioned graph data division method is within the preset capacity range, for example, controlled at the KB-MB level, the asynchronous IO library can be used to randomly read all the updated subgraphs in the partition from the partition of the disk, and the read updated subgraphs are transferred to the memory buffer, or the read updated subgraphs are directly transferred to the memory execution area for processing, so as to realize parallel processing of graph processing and IO operations. If the read updated subgraph is transferred to the memory buffer, the memory buffer also needs to be processed. The memory capacity occupied by the update subgraph is monitored, and when the memory capacity occupied reaches a preset threshold, the update subgraph in the memory buffer is transferred to the memory execution area for processing; during the running period, according to the preset graph algorithm, the update subgraphs in the memory execution area are processed in parallel, and for any update subgraph, according to the processing process, the state of each node in the update subgraph is updated, and according to the state of all nodes in the update subgraph, the state of the update subgraph is updated; for any update subgraph, if the state of the update subgraph reaches a preset termination state or the state of the update subgraph cannot continue to change, the processing of the update subgraph is stopped to obtain the processed update subgraph; during the writing period, the processed update subgraph is written to the disk.
[0055] For example, see Figure 7 , is a schematic diagram of a system architecture for processing graph data in a single-machine multi-core environment provided by the fourth embodiment of the application. Figure 7 As shown, the system includes a reading module, an executor, a writing module, a disk and a control module. The control module includes a scheduler, a state manager and a configuration manager. The system memory is divided into two parts: a buffer and a runtime, namely a memory buffer and a memory execution area. All updated subgraph partitions obtained by dividing the target graph data based on the above-mentioned graph data partitioning method are stored in the disk.
[0056] The system also provides application programming interfaces (APIs). Before processing graph data, you can set the preset graph algorithm for processing the updated subgraph based on the API interface, and set the algorithm parameters and parallel parameters based on the configuration manager. The overall process of graph data processing based on this system architecture can be: In a round of iterative processing, the reading module reads all the updated subgraphs in the partition from the partition of the disk, and transfers all the updated subgraphs in the partition read to the memory buffer; the scheduler monitors the memory capacity occupied by the updated subgraphs in the memory buffer, and when the preset threshold is reached, the updated subgraphs in the memory buffer are transferred to the memory execution area; the executor processes the updated subgraphs in the memory execution area in parallel according to the preset graph algorithm; for any updated subgraph, the state manager updates the state of each node in the updated subgraph according to the processing process, and updates the state of the updated subgraph according to the state of all nodes in the updated subgraph; for any updated subgraph, if the state of the updated subgraph reaches the preset termination state, the processing of the updated subgraph is stopped to obtain the processed updated subgraph, and the writing module writes the processed updated subgraph to the disk.
[0057] In the embodiment of the present application, when graph processing is performed in parallel on the update subgraphs obtained by the above division based on a preset graph algorithm, since good connectivity is maintained between the update subgraphs, the information synchronization speed between the update subgraphs is greatly promoted. During the processing, information can be transmitted more quickly between the update subgraphs, which improves the convergence efficiency of the graph processing and thus improves the performance of the graph processing.
[0058] See also Figure 8 , is a flow chart of a graph data processing method provided in Embodiment 5 of the present application, such as Figure 8 As shown, the state of each node includes an activated state and an inactivated state. The activated state is used to indicate that the node needs to be scheduled and processed in the current round, and the inactivated state is used to indicate that the node does not need to be scheduled and processed in the current round. In the above step S604, for any updated subgraph, if the state of the updated subgraph reaches a preset termination state, the processing of the updated subgraph is stopped, and the processed updated subgraph is obtained, which may include the following steps: Step S801, for any updated subgraph, determining the state of each node in the updated subgraph; Step S802: If the state of each node in the update subgraph is an inactive state, it is determined that the state of the update subgraph reaches a preset termination state, and the processing of the update subgraph is stopped to obtain a processed update subgraph.
[0059] In this embodiment, the state of each node includes an activated state and an inactivated state. The activated state is used to indicate that the node needs to be scheduled in the current round, and the inactivated state is used to indicate that the node does not need to be scheduled in the current round.
[0060] For any update subgraph, when the state of each node in the update subgraph is in an inactive state, that is, there is no need to continue processing any node in the update subgraph in the current round, that is, there is no need to continue processing the update subgraph, then it is determined that the state of the update subgraph has reached a preset termination state, the processing of the update subgraph is stopped, and the processed update subgraph is obtained.
[0061] In an embodiment of the present application, when the state of each node in the update subgraph is in an inactive state, it is determined that the state of the update subgraph has reached a preset termination state, and processing of the update subgraph is stopped to obtain a processed update subgraph, thereby avoiding unnecessary calculations and iterations and improving the efficiency of graph processing.
[0062] Corresponding to the graph data partitioning method in the above embodiment, Fig. 9 The structure block diagram of the graph data partitioning device provided in the sixth embodiment of the present application is shown. The graph data partitioning device is applied to Figure 1 For ease of description, only the parts related to the embodiment of the present application are shown.
[0063] See also Fig. 9 , the graph data partitioning device comprises: An acquisition module 91 is used to acquire target graph data, determine the out-degree of each node in the target graph data, and determine the node with the largest out-degree as the target node from the target graph data; A partitioning module 92 is used to partition the target graph data according to the target node to obtain at least two initial subgraphs, and determine edge nodes in each initial subgraph, wherein each initial subgraph includes the target node, and the edge node refers to a node included in at least two initial subgraphs; An adjustment module 93 is used to obtain the number of edge nodes in all initial subgraphs, and to adjust the node paths corresponding to the edge nodes in at least one initial subgraph with the number of edge nodes being within a preset value range as a goal, so as to obtain an updated subgraph; The first loop module 94 is used to, for any updated subgraph, if the storage capacity required to store the updated subgraph is not within a preset capacity range, then the updated subgraph is used as the target graph data, and the step of determining the out-degree of each node in the target graph data is returned to execute until all the updated subgraphs are obtained, wherein the updated subgraphs are used for partitioned storage in the disk.
[0064] Optionally, the adjustment module 93 includes: A first determining unit is configured to determine, for any initial subgraph, an edge node in the initial subgraph, and, for any edge node in the initial subgraph except the target node, determine a node path in the initial subgraph that includes the edge node; A first updating unit is configured to add, for any node path, the node path to any initial subgraph including the edge node except the initial subgraph, to obtain a subgraph to be evaluated; A calculation unit, used to determine the total number of nodes in each subgraph to be evaluated and the number of edge nodes in all subgraphs to be evaluated, and for any subgraph to be evaluated, respectively calculate the difference between the total number of nodes in the subgraph to be evaluated and the total number of nodes in each subgraph to be evaluated except the subgraph to be evaluated; The second determining unit is configured to determine that each subgraph to be evaluated is an updated subgraph if the calculated differences corresponding to all subgraphs to be evaluated are smaller than a preset value and the number of edge nodes in all subgraphs to be evaluated is within the preset value range.
[0065] Optionally, the adjustment module 93 further includes: The second loop unit is used to, if there is a subgraph to be evaluated whose corresponding difference is not less than the preset value, or the number of edge nodes in all subgraphs to be evaluated is not within the preset value range, then, for any subgraph to be evaluated, the subgraph to be evaluated is set as the initial subgraph, and return to execute the step of determining the edge nodes in the initial subgraph until each subgraph to be evaluated is determined to be an updated subgraph.
[0066] It should be noted that the information interaction, execution process and other contents between the above-mentioned modules are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0067] Corresponding to the graph data processing method of the above embodiment, Fig.10 The structural block diagram of the graph data processing device provided in the seventh embodiment of the present application is shown. For the convenience of explanation, only the part related to the embodiment of the present application is shown.
[0068] See also Fig.10 After all updated subgraphs are obtained by using the graph data partitioning method, the graph data partitioning device includes: The reading module 1001 is used to read all updated subgraphs in the partition from the partition of the disk in a round of iterative processing, and transfer all updated subgraphs in the partition read to the memory buffer; The scheduling module 1002 is used to monitor the memory capacity occupied by the updated subgraph in the memory buffer, and when the memory capacity reaches a preset threshold, transfer the updated subgraph in the memory buffer to the memory execution area; An operation module 1003 is used to perform parallel processing on the update subgraphs in the memory execution area according to a preset graph algorithm, and update the state of each update subgraph according to the processing process; The writing module 1004 is used for stopping processing any updated subgraph if the state of the updated subgraph reaches a preset termination state, obtaining a processed updated subgraph, and writing the processed updated subgraph into the disk.
[0069] Optionally, the running module 1003 includes: A node status management unit, configured to update the status of each node in any updated subgraph according to the processing procedure; A graph state management unit is used to update the state of the update subgraph according to the states of all nodes in the update subgraph.
[0070] Optionally, the writing module 1004 includes: A third determining unit, configured to determine, for any updated subgraph, a state of each node in the updated subgraph; The fourth determining unit is used to determine that the state of the updated subgraph reaches the preset termination state if the state of each node in the updated subgraph is the inactive state, stop processing the updated subgraph, and obtain the processed updated subgraph.
[0071] It should be noted that the information interaction, execution process and other contents between the above-mentioned modules are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0072] Fig.11 This is a schematic diagram of the structure of a computer device provided in Example 8 of the present application. Fig.11 As shown, the computer device of this embodiment includes: at least one processor ( Fig.11 Only one is shown), a memory, and a computer program stored in the memory and executable on at least one processor, wherein when the processor executes the computer program, the steps in any of the above-mentioned graph data partitioning method and graph data processing method embodiments are implemented.
[0073] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that Fig.11 These are merely examples of computer devices and do not constitute limitations on the computer devices. The computer devices may include more or fewer components than those shown in the figure, or a combination of certain components, or different components. For example, they may also include a network interface, a display screen, and an input device.
[0074] The processor may be a CPU, or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0075] The memory includes a readable storage medium, an internal memory, etc., wherein the internal memory may be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium may be a hard disk of a computer device, and in other embodiments, it may also be an external storage device of the computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Further, the memory may also include both an internal storage unit of the computer device and an external storage device. The memory is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of a computer program, etc. The memory may also be used to temporarily store data that has been output or is to be output.
[0076] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned method embodiment can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device capable of carrying computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0077] The present application implements all or part of the processes in the above-mentioned embodiment method, and may also be completed through a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiment when executing the computer program product.
[0078] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0079] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0080] In the embodiments provided in the present application, it should be understood that the disclosed devices / computer equipment and methods can be implemented in other ways. For example, the device / computer equipment embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0081] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0082] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A graph data partitioning method, characterized in that: The graph data partitioning method comprises: Obtain target graph data, determine the out-degree of each node in the target graph data, and determine the node with the largest out-degree from the target graph data as the target node; According to the target node, the target graph data is divided to obtain at least two initial subgraphs, and edge nodes in each initial subgraph are determined, wherein each initial subgraph includes the target node, and the edge node refers to a node included in at least two initial subgraphs; Obtaining the number of edge nodes in all initial subgraphs, adjusting the node paths corresponding to the edge nodes in at least one initial subgraph with the goal of ensuring that the number of edge nodes is within a preset value range, and obtaining an updated subgraph; For any updated subgraph, if the storage capacity required to store the updated subgraph is not within the preset capacity range, the updated subgraph is replaced by the target graph data, and the step of determining the out-degree of each node in the target graph data is returned to execute until all updated subgraphs are obtained, wherein the updated subgraph is used for partitioned storage on the disk.
2. The graph data partitioning method according to claim 1, characterized in that: The step of adjusting the node paths corresponding to the edge nodes in at least one initial subgraph with the number of the edge nodes being within a preset value range as a goal to obtain an updated subgraph includes: For any initial subgraph, determine the edge nodes in the initial subgraph, and for any edge node in the initial subgraph except the target node, determine the node path including the edge node in the initial subgraph; For any node path, add the node path to any initial subgraph including the edge node except the initial subgraph to obtain a subgraph to be evaluated; Determine the total number of nodes in each subgraph to be evaluated and the number of edge nodes in all subgraphs to be evaluated, and for any subgraph to be evaluated, respectively calculate the difference between the total number of nodes in the subgraph to be evaluated and the total number of nodes in each subgraph to be evaluated except the subgraph to be evaluated; If the calculated differences corresponding to all subgraphs to be evaluated are smaller than a preset value, and the number of edge nodes in all subgraphs to be evaluated is within the preset value range, each subgraph to be evaluated is determined to be an updated subgraph.
3. The graph data partitioning method according to claim 2, characterized in that: After respectively calculating the difference between the total number of nodes in the subgraph to be evaluated and the total number of nodes in each subgraph to be evaluated except the subgraph to be evaluated, the method further includes: If there is a subgraph to be evaluated whose corresponding difference is not less than the preset value, or the number of edge nodes in all subgraphs to be evaluated is not within the preset value range, then for any subgraph to be evaluated, the subgraph to be evaluated is set as the initial subgraph, and the step of determining the edge nodes in the initial subgraph is returned to execute until each subgraph to be evaluated is determined to be an updated subgraph.
4. A graph data processing method, characterized in that: After all updated subgraphs are obtained by using the graph data partitioning method according to any one of claims 1 to 3, the graph data processing method comprises: In one round of iterative processing, all updated subgraphs in the partition are read from the partition of the disk, and all updated subgraphs in the partition are transferred to the memory buffer; Monitoring the memory capacity occupied by the updated subgraph in the memory buffer, and when the memory capacity reaches a preset threshold, transmitting the updated subgraph in the memory buffer to the memory execution area; According to a preset graph algorithm, the update subgraphs in the memory execution area are processed in parallel, and according to the processing process, the state of each update subgraph is updated; For any updated subgraph, if the state of the updated subgraph reaches a preset termination state, the processing of the updated subgraph is stopped to obtain a processed updated subgraph, and the processed updated subgraph is written to the disk.
5. The graph data processing method according to claim 4, characterized in that: The step of updating the state of each updated subgraph according to the processing procedure includes: For any updated subgraph, according to the processing procedure, the state of each node in the updated subgraph is updated; The state of the update subgraph is updated according to the states of all nodes in the update subgraph.
6. The graph data processing method according to claim 5, characterized in that: The state of each node includes an activated state and an inactivated state, wherein the activated state is used to indicate that the node needs to be scheduled for processing in the current round, and the inactivated state is used to indicate that the node does not need to be scheduled for processing in the current round. For any updated subgraph, if the state of the updated subgraph reaches a preset termination state, the processing of the updated subgraph is stopped, and the processed updated subgraph is obtained, including: For any updated subgraph, determining the state of each node in the updated subgraph; If the state of each node in the update subgraph is the inactive state, it is determined that the state of the update subgraph reaches the preset termination state, and the processing of the update subgraph is stopped to obtain the processed update subgraph.
7. A graph data partitioning device, characterized in that: The graph data partitioning device comprises: An acquisition module is used to acquire target graph data, determine the out-degree of each node in the target graph data, and determine the node with the largest out-degree as the target node from the target graph data; A partitioning module is used to partition the target graph data according to the target node to obtain at least two initial subgraphs, and determine the edge nodes in each initial subgraph, wherein each initial subgraph includes the target node, and the edge nodes refer to nodes included in at least two initial subgraphs; An adjustment module is used to obtain the node number of edge nodes in all initial subgraphs, and to adjust the node path corresponding to the edge node in at least one initial subgraph with the node number of the edge node being within a preset value range as a goal, so as to obtain an updated subgraph; The first loop module is used to, for any updated subgraph, if the storage capacity required to store the updated subgraph is not within a preset capacity range, then the updated subgraph is used as the target graph data, and the step of determining the out-degree of each node in the target graph data is returned to execute until all the updated subgraphs are obtained, wherein the updated subgraphs are used for partitioned storage in the disk.
8. A graph data processing device, characterized in that: After all updated subgraphs are obtained by using the graph data partitioning method according to any one of claims 1 to 3, the graph data processing device comprises: A reading module, used for reading all updated subgraphs in the partition from the partition of the disk in a round of iterative processing, and transferring all updated subgraphs in the partition read to a memory buffer; A scheduling module, used to monitor the memory capacity occupied by the updated subgraph in the memory buffer, and when the memory capacity reaches a preset threshold, transfer the updated subgraph in the memory buffer to the memory execution area; An operation module, configured to perform parallel processing on the update subgraphs in the memory execution area according to a preset graph algorithm, and update the state of each update subgraph according to the processing process; The writing module is used for stopping processing any updated subgraph if the state of the updated subgraph reaches a preset termination state, obtaining a processed updated subgraph, and writing the processed updated subgraph into the disk.
9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the graph data partitioning method as described in any one of claims 1 to 3, or the graph data processing method as described in any one of claims 4 to 6.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the graph data partitioning method as described in any one of claims 1 to 3, or the graph data processing method as described in any one of claims 4 to 6.