Large-scale heterogeneous graph data mining method and device based on distributed shared memory

Through the distributed shared memory architecture and task delegation mechanism, the problems of low computing efficiency and communication overhead in large-scale graph data mining on heterogeneous devices are solved, efficient graph data mining is realized, and system performance and scalability are improved.

CN120386808APending Publication Date: 2025-07-29INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510437487.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art has low computing efficiency and high communication overhead when processing large-scale graph data on heterogeneous devices, making it difficult to effectively scale.

Method used

The distributed shared memory architecture is adopted, and the search tasks are sent to the data device for execution through the task delegation mechanism. Combined with task context management and message aggregation, it reduces duplicate data pull and traffic, and improves communication efficiency.

Benefits of technology

It significantly reduces the traffic by two orders of magnitude, improves the parallel efficiency of heterogeneous devices, improves the average performance by an order of magnitude, and can process larger-scale graph data and scale to multiple nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386808A_ABST
    Figure CN120386808A_ABST
Patent Text Reader

Abstract

The invention provides a large-scale heterogeneous graph data mining method and device based on a distributed shared memory. The large-scale heterogeneous graph data mining method comprises the steps that a distributed computing network used for graph data mining is obtained; graph data and a query graph are obtained, a task delegator distributes sub-graphs obtained by splitting the graph data to all heterogeneous computing devices in the distributed computing network, and search tasks are distributed to all the heterogeneous computing devices; an executor of the heterogeneous computing equipment executes the search task according to the local sub-graph and the query graph, and identifies whether the search task can be completed locally or not in the execution process, if yes, the search task is completed locally, and the isomorphic sub-graph obtained through execution is stored locally; otherwise, the task delegator marks the search task; and when all the search tasks of the distributed computing network are executed, gathering the local isomorphic sub-graphs of all the heterogeneous computing devices to obtain a graph data mining result of the query graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large-scale graph computing and graph data mining, and particularly relates to a method, device, electronic device, computer-readable storage medium, and computer program product for large-scale heterogeneous graph data mining based on distributed shared memory. Background Art

[0002] With the advent of the big data era, the demand for data processing and analysis has been increasing day by day. Graph computing can effectively process complex data relationships and is one of the key technologies for in-depth analysis of massive data. As an important part of the graph computing field, the core task of graph data mining is to find all subgraphs isomorphic to the query graph in the data graph. The huge search space makes it difficult to solve the entire graph computing task within polynomial time. To accelerate the search, existing work attempts to parallelize the graph data mining with heterogeneous devices (such as NVIDIA GPUs).

[0003] Graph data mining can be used for social network analysis to find a specified user group in a social network. At this time, the nodes in the graph data represent users, the edges represent a certain user behavior, such as following / blocking, and the pattern graph represents the user relationship to be found. Or graph data mining can be used for logic circuit synthesis to optimize a specific circuit in a circuit diagram into a circuit module with a smaller area, lower delay, and lower power consumption. At this time, the points in the graph data represent circuit elements (e.g., AND gates / NOT gates), and the pattern graph represents the circuit module that can be optimized.

[0004] However, the memory capacity of a single heterogeneous device is limited, making it difficult to scale the data size processed by the graph computing task. To process large-scale graph data, existing work attempts to partition the graph data onto multiple heterogeneous devices for parallel computing. This causes the neighbor lists of the vertices in the graph to be distributed on different heterogeneous devices. Such a graph partitioning method will cause the neighbor lists to be repeatedly fetched during the search process, resulting in a large amount of communication overhead.

[0005] To avoid expensive communication overhead, some work attempts to back up the graph data on multiple devices for task-level parallelism, or attempts to partition the graph data into multiple overlapping subgraphs (bin view) that store k-hop neighbors so that all the data required for the search can be included, enabling each heterogeneous device to process different subgraphs. In addition, some existing technologies attempt to use external memory to expand the graph data size, and when the heterogeneous device executes the graph data mining task, the required neighbor data can be obtained through PCIe. In a distributed homogeneous scenario, a message communication-based scheme can be adopted to process the graph computing task, and the communication overhead can be reduced by caching the neighbor data of high-degree vertices on each device.

[0006] In summary, existing work attempts to accelerate graph data mining using heterogeneous devices. However, due to the limited memory of heterogeneous devices, processing large-scale graph data is currently difficult. In distributed scenarios, graph data needs to be partitioned. Existing work focuses on more efficient graph partitioning methods or reducing communication overhead to handle large-scale graph data mining.

[0007] To handle distributed graph data mining, existing technologies are mainly divided into three types: data backup solutions (replication), vertical expansion (scaling up), and horizontal expansion (scaling out).

[0008] Data backup solutions store the entire graph on each heterogeneous device, completely eliminating communication. The simplest data backup solution is workload parallelism. Before starting the computation, the system divides the task into N equal parts (corresponding to the number of heterogeneous devices). Because each device stores the entire graph, each device can independently execute the mining task. However, this solution requires a large amount of device memory, limiting the scale of the graph data.

[0009] In a vertical scale-out solution, the system partitions graph data using host memory and transfers it to heterogeneous device memory in a pipelined manner for computation, alleviating the memory limitations of heterogeneous devices. To avoid communication, all data required for the search must be stored in each subgraph. However, this solution stores duplicate subgraph data, resulting in redundant storage overhead. Furthermore, this solution requires frequent PCIe data transfers, making vertical scale-out difficult to scale to heterogeneous devices.

[0010] The horizontal scaling solution involves partitioning the graph data into non-overlapping subgraphs. Each compute node independently processes its own subgraph data, communicating with other nodes via message passing. However, since each compute node cannot cache all neighbor data, the graph data mining process requires repeated remote data fetching, resulting in expensive communication overhead. When processing large graphs, communication volume and time become system performance bottlenecks. Summary of the Invention

[0011] The purpose of the present invention is to solve the problems of low computing efficiency of heterogeneous devices and high communication overhead in distributed scenarios in the above-mentioned prior art, and proposes a distributed graph mining optimization solution based on distributed shared memory. The task delegation mechanism is used to send tasks to the device where the data is located, avoiding repeated duplicate data pulling, thereby realizing efficient large-scale heterogeneous graph data mining.

[0012] In view of the shortcomings of existing technologies, such as Figure 9As shown in the figure, the present invention proposes a large-scale heterogeneous graph data mining method based on distributed shared memory, which includes:

[0013] Initial step: Obtain the distributed computing network for graph data mining, which includes multiple heterogeneous computing devices; construct a task dispatcher for distributing search tasks in the distributed computing network, construct a context manager for uniformly managing the memory of all the heterogeneous computing devices, and construct an executor for graph data mining within a single heterogeneous computing device;

[0014] Distribution step: Obtain graph data and a query graph. The task dispatcher distributes the subgraphs obtained by splitting the graph data to each heterogeneous computing device in the distributed computing network and assigns search tasks to each heterogeneous computing device;

[0015] Execution step: The executor of the heterogeneous computing device executes the search task according to the local subgraph and the query graph, and identifies whether the search task can be completed locally during the execution. If so, the search task is completed locally, and the isomorphic subgraph obtained by the execution is saved locally; otherwise, the task dispatcher marks the search task;

[0016] Re - distribution step: The context manager writes the marked search tasks into the specified heterogeneous computing device through the distributed shared memory of the distributed computing network, and executes the execution step again;

[0017] Summarization step: When all the search tasks of the distributed computing network have been executed, collect the isomorphic subgraphs local to all the heterogeneous computing devices to obtain the graph data mining result of the query graph.

[0018] In the large-scale heterogeneous graph data mining method based on distributed shared memory, the search task is to find all isomorphic subgraphs isomorphic to the query graph in the subgraph;

[0019] The subgraphs are partitioned onto different heterogeneous computing devices; each heterogeneous computing device maintains its local search state, saves the query results in the previous iterations and forms a prefix tree; in each iteration, the heterogeneous computing device reads the matching result P and the specified neighbor data from the prefix tree, and performs set operations to generate a new candidate set; after the current iteration is completed, the current search state is recorded and the prefix tree is updated; in a distributed scenario, if the data is not on the current heterogeneous computing device, the distributed shared memory is used to achieve data transmission between heterogeneous computing devices.

[0020] The described large-scale heterogeneous graph data mining method based on distributed shared memory is characterized in that when communicating between heterogeneous computing devices, each heterogeneous computing device first performs message communication within the node and sends it to the master device within the node, then the master devices between different nodes perform message passing, and finally the master device distributes the messages within the node.

[0021] The described large-scale heterogeneous graph data mining method based on distributed shared memory, wherein the reallocation step includes:

[0022] The context manager component preserves the search state of the search task with the tag, serializes the corresponding context, inserts it into the task queue of the remote heterogeneous computing device through the distributed shared memory, and the remote device resumes the search state through the context to continue executing the search task.

[0023] As Figure 10 shown, the present invention also proposes a large-scale heterogeneous graph data mining device based on distributed shared memory, which includes:

[0024] An initial module that obtains a distributed computing network for graph data mining, the distributed computing network includes multiple heterogeneous computing devices; constructs a task delegator for distributing search tasks in the distributed computing network, constructs a context manager for uniformly managing the memory of all the heterogeneous computing devices, and constructs an executor for graph data mining within a single heterogeneous computing device;

[0025] A distribution module that obtains graph data and a query graph, and the task delegator distributes the subgraphs obtained by splitting the graph data to each heterogeneous computing device in the distributed computing network and assigns search tasks to each heterogeneous computing device;

[0026] An execution module, where the executor of the heterogeneous computing device executes the search task according to the local subgraph and the query graph, and identifies whether the search task can be completed locally during the execution process. If so, the search task is completed locally, and the isomorphic subgraph obtained by the execution is saved locally; otherwise, the task delegator marks the search task;

[0027] A reallocation module, where the context manager writes the marked search task to the specified heterogeneous computing device through the distributed shared memory of the distributed computing network and executes the execution module again;

[0028] A summarization module, when all the search tasks of the distributed computing network have been executed, aggregates the isomorphic subgraphs local to all heterogeneous computing devices to obtain the graph data mining result of the query graph.

[0029] The described large-scale heterogeneous graph data mining device based on distributed shared memory, wherein the search task is to find all isomorphic subgraphs isomorphic to the query graph in the subgraph;

[0030] Sub - graphs are partitioned onto different heterogeneous computing devices; each heterogeneous computing device maintains its local search state, saves the query results in the previous several rounds of iterations and forms a prefix tree; in each round of iteration, the heterogeneous computing device reads the matching result P and the specified neighbor data from the prefix tree, and performs set operations to generate a new candidate set; after the current iteration is completed, the current search state is recorded and the prefix tree is updated; in a distributed scenario, if the data is not on the current heterogeneous computing device, the distributed shared memory is used to implement data transmission between heterogeneous computing devices.

[0031] For the large - scale heterogeneous graph data mining device based on distributed shared memory, when communicating between heterogeneous computing devices, each heterogeneous computing device first performs in - node message communication and sends it to the main device within the node, then the main devices between different nodes perform message passing, and finally the main device distributes the messages within the node;

[0032] The re - allocation module includes:

[0033] The context manager component retains the search state of the search task with the tag, serializes the corresponding context, inserts it into the task queue of the remote heterogeneous computing device through the distributed shared memory, and the remote device restores the search state through the context to continue executing the search task.

[0034] The present invention also proposes an electronic device, which includes the large - scale heterogeneous graph data mining device based on distributed shared memory described above. The electronic device is either connected to an information display device, and the information display device is used to display the graph data mining result with the display parameters, attributes set by the user or through an artificial intelligence model.

[0035] The present invention also proposes a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the large - scale heterogeneous graph data mining method based on distributed shared memory are implemented.

[0036] The present invention also proposes a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the large - scale heterogeneous graph data mining method based on distributed shared memory are implemented.

[0037] As can be seen from the above solutions, the advantages of the present invention are:

[0038] This invention focuses on the time-consuming communication problem in distributed graph mining, aiming to reduce the communication volume while improving the communication efficiency: (1) Through task delegation, this invention transfers the search tasks to heterogeneous devices where the graph data is located for execution, rather than pulling the required data to the local for execution, thus reducing the communication volume by two orders of magnitude. (2) Since distributed shared memory avoids explicit message synchronization of the CPU, a large number of small messages will be generated during the computing process by heterogeneous devices. This invention improves the bandwidth utilization through message aggregation. The results show that, compared with the state-of-the-art systems, the average performance of this invention has increased by one order of magnitude. With the same memory resources, it can process larger-scale graph data and can be extended to multiple nodes, with a parallel efficiency of about 66%. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 FIG. for a communication-friendly distributed graph data mining solution

[0040] Figure 2 FIG. for the flowchart of graph data mining with distributed shared memory

[0041] Figure 3 FIG. for the message characteristics in distributed graph data

[0042] Figure 4 FIG. for the grouped communication mode within a node

[0043] Figure 5 FIG. for the multi-stage communication mode between nodes

[0044] Figure 6 FIG. for the comparison of message passing mode and task delegation mode

[0045] Figure 7 FIG. for the effect diagram of message volume reduction in task delegation mode under real scenarios

[0046] Figure 8 FIG. for task context description and flowchart

[0047] Figure 9 FIG. for the flowchart of the method of this invention

[0048] Figure 10 FIG. for the device module diagram of this invention

[0049] Figure 11 FIG. for the schematic structural diagram of the first electronic device of this invention

[0050] Figure 12 FIG. for the schematic structural diagram of the application environment of the first electronic device of this invention

[0051] Figure 13 FIG. for the schematic structural diagram of the second electronic device of this invention

[0052] Reference Signs:

[0053] A - The first electronic device;

[0054] B - Large-scale heterogeneous graph data mining device based on distributed shared memory;

[0055] C - Data acquisition device;

[0056] D - Information display device;

[0057] 1000 - The second electronic device;

[0058] Ⅰ - Computing unit;

[0059] Ⅱ - ROM;

[0060] Ⅲ - RAM;

[0061] Ⅳ - Bus;

[0062] Ⅴ - Interface;

[0063] Ⅵ - Input unit;

[0064] Ⅶ - Output unit;

[0065] Ⅷ - Storage medium;

[0066] Ⅸ - Communication unit. Detailed Implementation Manner

[0067] It should be noted that in this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0068] Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the said element.

[0069] The processor described in the present invention is the control center of an electronic device, which can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0070] Optionally, the processor can execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.

[0071] In a specific implementation, as an embodiment, the processor may include one or more CPUs. Each of these processors can be a single-CPU or a multi-CPU. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). The electronic device may include: servers, desktop computers, laptop computers, smartphones, tablet computers, embedded computers, etc., where the embedded computer includes vehicles and robots, etc.

[0072] The memory is used to store the software program for implementing the solution of the present invention and is controlled by the processor for execution. The specific implementation method can refer to the above method embodiments and will not be elaborated here.

[0073] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not limit it. The actual knowledge structure recognition device may include more or fewer components than shown in the drawings, or combine certain components, or have a different component arrangement.

[0074] The above embodiments can be implemented in whole or in part through software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired method (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0075] It should also be understood that the term "and / or" in this document simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " in this document generally indicates an "or" relationship between the related objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0076] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0077] It should also be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0078] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of the device or unit, which can be electrical, mechanical or other forms.

[0079] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0080] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0081] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0082] The present invention analyzes the existing distributed graph data mining technology and finds that the existing technology has problems such as low computing efficiency and limited scalability of heterogeneous devices when processing large-scale graph data. These systems usually require a CPU to perform graph partitioning and sub-graph data management. They rely on the CPU for computational synchronization and frequently exchange sub-graph data between the main memory and the video memory of heterogeneous devices, which seriously hinders the computing efficiency of heterogeneous devices. The present invention abstracts the entire heterogeneous device cluster into a single heterogeneous device with a huge shared memory capacity. This solution helps to implement non-overlapping graph partitioning schemes in heterogeneous device clusters, so that the remote data required for heterogeneous computing can be obtained on demand through high-speed interconnection. This solution can effectively eliminate the computing interruption of heterogeneous devices caused by data transmission initiated by the CPU.

[0083] To achieve this goal, this paper proposes a novel distributed optimization solution for heterogeneous graph data mining. This solution centrally manages the memory of an entire heterogeneous cluster, supporting communication across heterogeneous devices. However, to address the high communication overhead associated with graph data mining, reducing overall system communication volume and improving communication efficiency are key challenges in this paper.

[0084] Driven by this goal, the present invention proposes Figure 1 The technical optimization route shown in the figure mainly considers three aspects:

[0085] The first is to design an execution engine (Executor), which is mainly responsible for graph data mining within a single heterogeneous device to ensure that the entire system can efficiently perform heterogeneous computing. The second is to design a delegation engine (Delegator), which packages search tasks and sends them to the heterogeneous device where the data is located. The communication mode is changed from transmitting data to the location of the task to sending tasks to the location of the data. At this time, the task refers to the task generated when the calculation is performed, which is generated by the context manager (ContextManager) by analyzing the read data. The third is to design a context management engine (ContextManager), which manages the interaction with the video memory of all heterogeneous devices and provides a low-level interface (low-levelAPI) to support the semantics of distributed shared memory.

[0086] The overall input data of the present invention includes a data graph and a pattern graph. When a heterogeneous device reads data from the local video memory, the execution engine is called to generate a new computing task. At this time, "reading data" means reading the topological information of the data graph from the video memory (that is, the connection relationship between points, usually stored in the CSR format), as well as the state information materialized during the calculation process. The delegation engine analyzes which remote heterogeneous devices these new tasks need to be sent to, and calls the context management engine to package these tasks and send them to the specified device at the same time. In this way, the entire system can avoid frequent data transmission, thereby significantly reducing the amount of communication and improving communication efficiency. Therefore, the present invention includes the following key technical points:

[0087] Key Point 1: A distributed shared memory graph data mining method. This method can efficiently manage heterogeneous memory in distributed graph mining, shielding communication between underlying heterogeneous devices.

[0088] Key point 2: A graph data mining communication strategy based on task delegation. This strategy sends tasks to the data location, avoiding repeated remote data communication and significantly reducing the overall communication volume of distributed graph mining.

[0089] Key point 3: Task context management strategy in distributed graph mining. Only the necessary data for the current round of calculation of a single graph mining task is retained, while multiple tasks are packaged and sent to remote heterogeneous devices, making the overall communication mode more regular and significantly improving communication efficiency.

[0090] To illustrate the above-mentioned features and effects of the present invention more clearly and easily, the following examples are provided and further described in detail with reference to the accompanying drawings. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are for illustrative purposes only. The scope of protection of the present invention is not limited to the disclosed embodiments; the present invention is defined by the appended claims.

[0091] This paper will explain the specific implementation methods of distributed heterogeneous graph data mining from three aspects: shared memory architecture design, task delegation, and task context management.

[0092] Shared memory architecture design:

[0093] The task of graph data mining is to find all subgraphs isomorphic to the query graph (pattern graph) within a large-scale graph dataset. Current graph mining algorithms employ a set-centric programming abstraction. The search process can be described as a nested loop: each level of the loop generates a new candidate set and search task as input to the next level. The search process terminates when all edges in the query graph are matched to edges in the data graph.

[0094] Figure 2The graph data mining workflow under our shared memory architecture is shown. The graph data is divided into different heterogeneous devices. Each heterogeneous device needs to maintain its own search state, save the query results in the previous iterations and form a prefix tree. In each iteration, the heterogeneous device reads the matching result P and the specified neighbor data from the prefix tree, and performs set operations to generate a new candidate set. After the current iteration is completed, the system will record the current search status and update the prefix tree. In a distributed scenario, if the data is not on the current heterogeneous device, the system will enable cross-device data transmission. Using distributed shared memory, data transmission can be simplified to the GET primitive, that is, specifying the data address to obtain remote neighbor data. The distributed shared memory architecture relies on high-bandwidth communication within and between nodes. However, in the process of graph data mining, a large number of irregular small messages will be generated on heterogeneous devices. For example Figure 3 As shown in the data mining scenario, an average of 93% of the message size is less than 200 bytes. And the frequency of these small messages is as high as 10 9 Due to the skew property of graph data, these message sizes vary significantly. This leads to even more irregular communication and memory access patterns. Furthermore, the high concurrency of heterogeneous devices causes different computing units to handle different computational tasks. This results in a large amount of random cross-device communication in a distributed shared memory architecture.

[0095] To achieve an efficient distributed shared memory architecture, we adopt different communication strategies within and between nodes, e.g. Figure 4 As shown. "Node" refers to the computing hardware in distributed computing, and "vertex" refers to a point in the graph data. Within the node, we divide the communication bandwidth equally into m parts (m is the number of heterogeneous devices in the node). Different cross-device requests are grouped according to the target device address. Remote communication is carried out independently in different groups. Under this strategy, each heterogeneous device can use 1 / (m-1) of the network bandwidth, so there will be no bandwidth competition. For cross-node communication, we adopt a multi-stage communication mode, such as Figure 5 As shown, any two heterogeneous devices do not communicate directly. Instead, each device first sends a message within the node to the master device within the node, then the master devices between different nodes carry out message transmission, and finally the master device distributes the message within the node.

[0096] Task delegation engine:

[0097] During the execution of graph data mining, a large number of computational subtasks (such as subtasks of search tasks) are generated. The core calculation of each subtask is to determine the connection relationship between two vertices. Traditional solutions usually adopt a communication model based on message passing, in which heterogeneous devices pull remote graph data on demand during each round of calculation. However, this solution requires repeatedly pulling neighbor data across devices, resulting in low communication efficiency. In actual calculations, we found that the size of each subtask is much smaller than the size of a single neighbor data. Leveraging this observation, we designed a communication model based on task delegation.

[0098] To simplify the description, Figure 6 In this paper, we present the workflow for executing search tasks on two heterogeneous devices. The subgraph consisting of vertices 0, 3, 5, 6 and their corresponding neighbors is saved on device 0, while the subgraph consisting of vertices 1, 2, 4, 7 and their corresponding neighbors is saved on device 1. In the message passing mode, device 0 needs to request the required graph data from device 1, that is, to pull the neighbor data of vertex 1 locally to determine whether there is a connection relationship between vertex 1 and vertex 7. In the task delegation scenario, device 0 sends vertex 1 to device 1, which then determines the connection relationship and returns the determination result to device 0. Under the task delegation scheme, the communication volume is reduced from |N(1)| = 6 to |1,7| = 2.

[0099] In real graph data mining tasks, the candidate set will be continuously reduced and filtered after multiple rounds of iterations, which makes the size of the subtask much smaller than the size of the neighbor data. Figure 7 In this paper, we analyzed the communication volume of message passing and task delegation schemes at different iteration rounds in two typical graph task scenarios. The results show that the message size in the task delegation scheme decreases significantly with each iteration round. After three iterations, the communication volume of task delegation is reduced to 1% of that of message passing.

[0100] Task context management engine:

[0101] In high-concurrency scenarios across heterogeneous devices, efficiently implementing task delegation is a significant challenge. Inspired by the process context switching mechanism in operating systems, we propose a task context management engine. Specifically, this component maintains context related to the current search state, serializes the required information, and inserts the context into the task queue of the remote heterogeneous device. When the remote device receives this message, it can use the context to restore the search state and resume graph computation tasks.

[0102] In graph data mining computations, the search task can be decomposed into multiple iterations (subtasks), each of which attempts to match a single edge in the pattern graph. Figure 8As shown, in this subtask, the system instantiates its context into three parts: the vertex sets S and P, and the operation op. Among them, S and P represent the candidate sets of the two vertices of the edge being matched, and the subtask needs to check which vertices in S are connected to the vertices in P. In each subtask, the system will enumerate all vertices v in S, retrieve its neighbors N(v), and then calculate The symbol represents a set of predefined set operations.

[0103] To serialize <S, P, op>, the task context management engine writes this data to contiguous device memory. After the context is serialized, it is directly written to the remote task queue. The engine adopts a producer - consumer computing model, and each heterogeneous device can read tasks from the local task queue, parse the context, decompose the atomic tasks, and independently execute subsequent calculations.

[0104] Figure 1 The working process of the present invention is also shown, which includes two workflows (as indicated by the arrows in the figure):

[0105] (1) Local workflow:

[0106] a) First step: The ContextManager reads data and status information from local memory to form tasks;

[0107] b) Second step: The ContextManager assigns the tasks to be processed to the Executor for execution;

[0108] c) Third step: The Executor returns the processing results to the ContextManager;

[0109] d) Fourth step: The ContextManager writes the output results back to the local GPU memory.

[0110] (2) Remote workflow

[0111] a) First step: The same as the local workflow;

[0112] b) Second step: The same as the local workflow;

[0113] c) Third step: When the Executor reads the task information and finds that the local machine does not hold the data required for the task and determines that the task cannot be completed locally, it hands it over to the Delegator;

[0114] d) Fourth step: The Delegator disassembles the task into sub - tasks that can be delegated according to the location of the data required for the task and hands it over to the ContextManager. In particular, if all the data required for the task is on the same computing node, there is no need to disassemble.

[0115] e) Step 5: The context manager writes the subtask status to the remote GPU’s video memory.

[0116] Access local workflows on remote GPUs.

[0117] When all the above processes are completed and no new tasks are generated, the entire workflow ends. At this time, all subgraphs that are isomorphic to the query graph (pattern graph) are stored in the GPU memory and output to the file.

[0118] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in conjunction with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0119] like Figure 10 As shown, the present invention also proposes a large-scale heterogeneous graph data mining device based on distributed shared memory, which includes:

[0120] An initialization module acquires a distributed computing network for graph data mining, the distributed computing network comprising multiple heterogeneous computing devices; constructs a task delegator for distributing search tasks in the distributed computing network, constructs a context manager for uniformly managing the memory of all the heterogeneous computing devices, and constructs an executor for graph data mining within a single heterogeneous computing device;

[0121] The distribution module obtains graph data and query graphs, and the task delegator distributes the subgraphs obtained by splitting the graph data to various heterogeneous computing devices in the distributed computing network and assigns search tasks to each heterogeneous computing device;

[0122] An execution module, wherein the executor of the heterogeneous computing device executes the search task according to the local subgraph and the query graph, and identifies during the execution process whether the search task can be completed locally. If so, the search task is completed locally and the homogeneous subgraph obtained by the execution is saved locally; otherwise, the task delegator marks the search task;

[0123] a redistribution module, wherein the context manager writes the marked search task to a designated heterogeneous computing device through the distributed shared memory of the distributed computing network and executes the execution module again;

[0124] The aggregation module, when all search tasks of the distributed computing network have been executed, aggregates the local isomorphic subgraphs of all heterogeneous computing devices to obtain the graph data mining results of the query graph.

[0125] The large-scale heterogeneous graph data mining device based on distributed shared memory, where the search task is to find all isomorphic subgraphs isomorphic to the query graph in the subgraph;

[0126] The subgraphs are partitioned onto different heterogeneous computing devices; each heterogeneous computing device maintains its local search state, saves the query results in the previous iterations and forms a prefix tree; in each iteration, the heterogeneous computing device reads the matching result P from the prefix tree and the specified neighbor data, and performs set operations to generate a new candidate set; after the current iteration is completed, the current search state is recorded and the prefix tree is updated; in a distributed scenario, if the data is not on the current heterogeneous computing device, the distributed shared memory is used to implement data transmission between heterogeneous computing devices.

[0127] The large-scale heterogeneous graph data mining device based on distributed shared memory, where during communication between heterogeneous computing devices, each heterogeneous computing device first performs intra-node message communication and sends it to the master device within the node, then the master devices between different nodes perform message passing, and finally the master device performs message distribution within the node;

[0128] The reallocation module includes:

[0129] The context manager component preserves the search state of the tagged search task, serializes the corresponding context and inserts it into the task queue of the remote heterogeneous computing device through the distributed shared memory, and the remote device resumes the search state through the context to continue executing the search task.

[0130] As Figure 11 shown, in another embodiment of the present invention, a first electronic device A is further proposed, which includes the large-scale heterogeneous graph data mining device based on distributed shared memory described above.

[0131] As Figure 12 shown, the first electronic device A can also be connected to the data acquisition device C and the information display device D through a wired or wireless information transmission scheme. The data acquisition device C is used to acquire graph data and query graphs. For example, when the usage scenario is video recommendation, the graph data is the entire user video viewing interaction relationship graph accumulated by the video platform, and the query graph is the user historical video viewing interaction relationship graph of the video to be recommended; the information display device D is used to display the graph data mining results analyzed by the present invention, such as recommended video content.

[0132] Among them, the information display device D can process the data output by the first electronic device A based on the information display mechanism to improve the readability of the data output by the first electronic device A. The information display mechanism can be preset manually. For example, the data output by the first electronic device A is visually displayed, and it can display according to the display parameters and / or attributes set by the user. The display parameters can be, for example, the display data range, and the display attributes can be, for example, the display font, color, whether to scroll and play, etc. Present the key information specified by the user to the user, enabling the user to understand this information more timely without having to access the secondary page or scroll the page, saving the user's operations. Or the information display mechanism can be an artificial intelligence AI display model, which can learn the key information of the user based on the user's previous usage habits, such as viewing duration, click times, editing times, etc., and then automatically present rich and necessary key information to the user.

[0133] The present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a readable storage medium. When the computer program is executed by a processor, the computer can execute the large-scale heterogeneous graph data mining method based on distributed shared memory provided by the above various methods.

[0134] The present invention also proposes a storage medium VIII in another embodiment for storing a computer program for executing the large-scale heterogeneous graph data mining method based on distributed shared memory. It should be understood that the storage medium in the embodiment of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM).

[0135] Figure 13 A schematic block diagram of a second electronic device 1000 that can be used to implement an embodiment of the present invention is shown. The second electronic device 1000 electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The second electronic device 1000 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein. The second electronic device 1000 may be the same as or different from the first electronic device A.

[0136] The second electronic device 1000 includes a computing unit I, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory II (ROM) or a computer program loaded from a storage medium VIII into a random access memory (RAM) III. Various programs and data required for the operation of the device 1000 can also be stored in the RAM III. The computing unit I, ROM II, and RAM III are connected to each other via a bus IV. An input / output (I / O) interface V is also connected to the bus IV.

[0137] Multiple components in the second electronic device 1000 are connected to the I / O interface V, including: an input unit VI, such as a keyboard and mouse; an output unit VII, such as various types of displays and speakers; a storage medium VIII, such as a magnetic disk and optical disk; and a communication unit IX, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit IX allows the second electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0138] Computing unit I can be various general and / or special processing components with processing and computing capabilities. Some examples of computing unit I include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. Computing unit I performs the various methods and processes described above, such as method steps S1-S5. For example, in some embodiments, the method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage medium VIII. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1000 via ROM II and / or communication unit IX. When the computer program is loaded into RAM III and executed by computing unit I, one or more steps of the method described above can be performed. Alternatively, in other embodiments, computing unit I can be configured to execute the method in any other appropriate manner (e.g., by means of firmware).

[0139] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.

Claims

1. A large-scale heterogeneous graph data mining method based on distributed shared memory, characterized in that, Including: An initial step of obtaining the distributed computing network for graph data mining, where the distributed computing network includes multiple heterogeneous computing devices; Constructing a task dispatcher for distributing search tasks in the distributed computing network, constructing a context manager for uniformly managing the memory of all the heterogeneous computing devices, and constructing an executor for graph data mining inside a single heterogeneous computing device; A distribution step of obtaining graph data and a query graph, where the task dispatcher distributes the subgraphs obtained by splitting the graph data to each heterogeneous computing device in the distributed computing network and assigns search tasks to each heterogeneous computing device; An execution step, where the executor of the heterogeneous computing device executes the search task according to the local subgraph and the query graph, and identifies during the execution whether the search task can be completed locally. If so, the search task is completed locally and the isomorphic subgraph obtained by the execution is saved locally; otherwise, the task dispatcher marks the search task; A reallocation step, where the context manager writes the marked search task into a specified heterogeneous computing device through the distributed shared memory of the distributed computing network and executes the execution step again; A summarization step, when all search tasks of the distributed computing network have been executed, aggregating the isomorphic subgraphs local to all heterogeneous computing devices to obtain the graph data mining result of the query graph.

2. The large-scale heterogeneous graph data mining method based on distributed shared memory according to claim 1, characterized in that The search task is to find all isomorphic subgraphs isomorphic to the query graph in the subgraph; The subgraphs are partitioned onto different heterogeneous computing devices; each heterogeneous computing device maintains its local search state, saves the query results in the previous iterations and forms a prefix tree; In each iteration, the heterogeneous computing device reads the matching result P and the specified neighbor data from the prefix tree and performs set operations to generate a new candidate set; After the current iteration is completed, the current search state is recorded and the prefix tree is updated; In a distributed scenario, if the data is not on the current heterogeneous computing device, the distributed shared memory is used to implement data transmission between heterogeneous computing devices.

3. The large-scale heterogeneous graph data mining method based on distributed shared memory according to claim 1, characterized in that When communicating between the heterogeneous computing devices, each heterogeneous computing device first performs intra-node message communication and sends it to the master device within the node, then the master devices between different nodes perform message passing, and finally the master device performs message distribution within the node.

4. The large-scale heterogeneous graph data mining method based on distributed shared memory according to claim 1, characterized in that The reallocation step includes: The context manager component preserves the search state of the marked search task, serializes the corresponding context and inserts it into the task queue of the remote heterogeneous computing device through the distributed shared memory, and the remote device resumes the search state through the context to continue executing the search task.

5. A large-scale heterogeneous graph data mining device based on distributed shared memory, characterized in that, Including: An initial module that obtains a distributed computing network for graph data mining, where the distributed computing network includes multiple heterogeneous computing devices; Constructing a task dispatcher for distributing search tasks in the distributed computing network, constructing a context manager for uniformly managing the memory of all the heterogeneous computing devices, and constructing an executor for graph data mining inside a single heterogeneous computing device; The distribution module obtains graph data and a query graph. The task dispatcher distributes the subgraphs obtained by splitting the graph data to each heterogeneous computing device in the distributed computing network and assigns search tasks to each heterogeneous computing device; The execution module. The executor of the heterogeneous computing device executes the search task according to the local subgraph and the query graph, and identifies whether the search task can be completed locally during the execution. If so, the search task is completed locally, and the isomorphic subgraph obtained by the execution is saved locally; otherwise, the task dispatcher marks the search task; The reallocation module. The context manager writes the marked search task into the specified heterogeneous computing device through the distributed shared memory of the distributed computing network and executes the execution module again; The aggregation module. When all search tasks of the distributed computing network have been executed, it aggregates the isomorphic subgraphs local to all heterogeneous computing devices to obtain the graph data mining result of the query graph.

6. The large-scale heterogeneous graph data mining device based on distributed shared memory according to claim 5, wherein, The search task is to find all isomorphic subgraphs isomorphic to the query graph in the subgraph; The subgraphs are partitioned onto different heterogeneous computing devices; each heterogeneous computing device maintains its local search state, saves the query results in the previous iterations and forms a prefix tree; In each iteration, the heterogeneous computing device reads the matching result P and the specified neighbor data from the prefix tree and performs set operations to generate a new candidate set; After the current iteration is completed, the current search state is recorded and the prefix tree is updated; In a distributed scenario, if the data is not on the current heterogeneous computing device, the distributed shared memory is used to achieve data transmission between heterogeneous computing devices.

7. The large-scale heterogeneous graph data mining device based on distributed shared memory according to claim 5, characterized in that, When communicating between the heterogeneous computing devices, each heterogeneous computing device first performs in-node message communication and sends it to the master device within the node, then the master devices between different nodes perform message passing, and finally the master device performs message distribution within the node; The reallocation module includes: The context manager component, which retains the search state of the marked search task, serializes the corresponding context and inserts it into the task queue of the remote heterogeneous computing device through the distributed shared memory. The remote device resumes the search state through the context and continues to execute the search task.

8. An electronic device, characterized in that, Including a large-scale heterogeneous graph data mining device based on distributed shared memory as described in claims 5-7. The electronic device is connected to an information display device, and the information display device is used to display the graph data mining result with the display parameters, attributes set by the user or through an artificial intelligence model.

9. A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the large-scale heterogeneous graph data mining method based on distributed shared memory as described in any one of claims 1-4.

10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the large-scale heterogeneous graph data mining method based on distributed shared memory as described in any one of claims 1-4.