Graph Data Mining Method, Device, Electronic Device and Machine-readable Storage Medium

By introducing a distributed memory management system into the data processing system of Apache Spark architecture, the problem of the separation of the big data processing framework and the graph deep learning framework is solved, and efficient graph data mining is achieved.

CN113867983BActive Publication Date: 2025-06-24HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111075298.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-14
Publication Date
2025-06-24
Estimated Expiration
2041-09-14

AI Technical Summary

Technical Problem

The existing big data processing framework is separated from the graph deep learning framework, resulting in efficient integration of data preprocessing and GNN model training becoming a difficult problem.

Method used

By introducing a distributed memory management system in a data processing system based on Apache Spark architecture, the graph structure data is divided into sub-graph data, and graph neural network model training is carried out in the distributed memory management system, and finally used for training and prediction of machine learning models.

Benefits of technology

It realizes efficient data interaction between the big data processing framework and the graph neural network framework, avoids frequent data transfer, and improves the execution efficiency of graph data mining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113867983B_ABST
    Figure CN113867983B_ABST
Patent Text Reader

Abstract

The present application provides a graph data mining method, apparatus, electronic device, and machine-readable storage medium. The method includes: preprocessing original data to obtain graph structure data; splitting the graph structure data according to the training strategy of a distributed graph neural network to obtain a plurality of sub-graph data, and storing the sub-graph data in a distributed memory management system; constructing a distributed graph neural network training function, and using the distributed graph neural network training function to perform distributed graph neural network model training according to the sub-graph data stored in the distributed memory management system, and storing the obtained Embedding in the distributed memory management system; training and predicting an ML model according to the Embedding saved in the distributed memory management system. This method can improve the execution efficiency of graph data mining.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to data processing technologies, and in particular, to a graph data mining method, apparatus, electronic device, and machine-readable storage medium. Background Art

[0002] Currently, for the processing of large amounts of data, big data processing frameworks represented by Apache Spark have emerged. They usually have good scalability and friendly programming interfaces, and can process large amounts of data quickly and effectively.

[0003] With the development of artificial intelligence technologies, deep learning methods have achieved great success in intelligent applications on Euclidean data such as images and texts. However, in reality, many data naturally belong to non-Euclidean graph structures, such as social networks, knowledge graphs, and molecular structures, etc. Drawing on the achievements of deep learning on Euclidean data, researchers have proposed various graph neural network models (Graph Neural Networks, abbreviated as GNN) for graph structures, which have been widely applied in fields such as search, recommendation, and drug research and development.

[0004] In the general process of the implementation of graph mining algorithms, people use big data processing frameworks to perform data preprocessing on massive raw data, then use graph neural network frameworks to obtain Embedding (representation), and finally return to the big data processing framework to complete the training and inference of machine learning (abbreviated as ML) algorithms.

[0005] However, the current big data processing frameworks and graph deep learning frameworks are disjointed, which makes it a problem to be solved to efficiently integrate data preprocessing and GNN model training. Summary of the Invention

[0006] In view of this, the present application provides a graph data mining method, apparatus, electronic device, and machine-readable storage medium.

[0007] Specifically, the present application is implemented through the following technical solutions:

[0008] According to a first aspect of an embodiment of the present application, a graph data mining method is provided, which is applied to a data processing system implemented based on the Apache Spark architecture. The method includes:

[0009] Preprocess the raw data to obtain graph-structured data;

[0010] According to the training strategy of the distributed graph neural network, split the graph-structured data to obtain multiple sub-graph data, and store the sub-graph data in a distributed memory management system;

[0011] Construct a distributed graph neural network training function. Using the distributed graph neural network training function, perform distributed graph neural network model training based on the sub-graph data stored in the distributed memory management system, and store the obtained Embedding in the distributed memory management system;

[0012] Train and predict the ML model based on the Embedding saved in the distributed memory management system.

[0013] According to the second aspect of the embodiments of the present application, there is provided a graph data mining device, which is applied to a data processing system implemented based on the Apache Spark architecture. The device includes:

[0014] A preprocessing unit for preprocessing the original data to obtain graph-structured data;

[0015] A splitting unit for splitting the graph-structured data into multiple sub-graph data according to the training strategy of the distributed graph neural network, and storing the sub-graph data in the distributed memory management system;

[0016] A training unit for constructing a distributed graph neural network training function. Using the distributed graph neural network training function, perform distributed graph neural network model training based on the sub-graph data stored in the distributed memory management system, and store the obtained Embedding in the distributed memory management system;

[0017] A mining unit for training and predicting the ML model based on the Embedding saved in the distributed memory management system.

[0018] According to the third aspect of the embodiments of the present application, there is provided an electronic device, including a processor and a machine-readable storage medium. The machine-readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor is used to execute the machine-executable instructions to implement the above-mentioned graph data mining method.

[0019] According to the fourth aspect of the embodiments of the present application, there is provided a machine-readable storage medium. The machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by a processor, the above-mentioned graph data mining method is implemented.

[0020] The graph data mining method of the embodiment of the present application, by combining the big data processing framework and the graph neural network framework and introducing a distributed memory management system of shared memory, opens up the data transmission between Spark and the graph neural network training architecture, efficiently realizes data fusion, avoids frequent data movement between the big data processing framework and the graph neural network framework, improves the data interaction efficiency between the big data framework and the graph neural network framework, and thus improves the execution efficiency of graph data mining. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a flowchart of a graph data mining method shown in an exemplary embodiment of the present application;

[0022] Figure 2 It is an architectural diagram of a data processing system implemented based on the Apache Spark architecture, shown in an exemplary embodiment of the present application;

[0023] Figure 3 It is a flowchart of a graph data mining method shown in an exemplary embodiment of the present application;

[0024] Figure 4 is a schematic diagram of a graph structure segmentation shown in an exemplary embodiment of the present application;

[0025] Figure 5 is a structural schematic diagram of a graph data mining device shown in an exemplary embodiment of the present application;

[0026] Figure 6 It is a schematic diagram of the hardware structure of an electronic device shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0027] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0028] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.

[0029] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, some technical terms involved in the embodiments of the present application are briefly explained below.

[0030] Apache Spark (abbreviation: Spark): A distributed big data processing framework, usually used for data preprocessing in graph data mining.

[0031] Spark application: A Spark job, an application program written by users, which is submitted to the Spark cluster for execution.

[0032] Horovod: A distributed deep learning training framework. Using this framework, it is relatively convenient to complete distributed training.

[0033] DGL (Deep Graph Library): A graph deep learning framework that can use multiple deep learning frameworks as the backend.

[0034] Apache Arrow: A memory-based columnar storage format that can be converted into multiple other data formats.

[0035] Embedding (representation): A low-dimensional vector representation. For example, Word Embedding is a low-dimensional vector representation of a word.

[0036] In order to make the above-mentioned objectives, features, and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0037] Please refer to Figure 1 , which is a schematic flowchart of a graph data mining method provided by an embodiment of the present application. Among them, this graph data mining method can be applied to a data processing system implemented based on the Apache Spark architecture, such as Figure 1 shown, this graph data mining method may include the following steps:

[0038] Step S100: Preprocess the original data to obtain graph structure data.

[0039] Step S110: According to the training strategy of the distributed graph neural network, split the graph structure data to obtain multiple sub-graph data, and store the multiple sub-graph data in the distributed memory management system.

[0040] In the embodiments of the present application, the original data can be read and preprocessed through Spark to obtain graph structure data.

[0041] Exemplarily, the graph structure data can be data required for graph construction, usually in the form of a binary tuple or a triple tuple, and each tuple represents an edge on the graph.

[0042] In an embodiment of the present application, Spark can be used to divide the graph structure data into sub-graphs so that in the subsequent process, multiple work nodes can be used to utilize a distributed graph neural network training method to perform graph neural network training based on each sub-graph data, thereby improving the efficiency of graph neural network training.

[0043] Exemplarily, when the graph structure data is sub-divided, it is necessary to divide the graph structure data according to the training strategy of the distributed graph neural network. For example, according to the requirements of the distributed graph neural network training for the graph deep learning framework (which may be referred to as the target graph deep learning framework) used for the pre-configured distributed graph neural network training, the graph structure data is divided.

[0044] For example, assuming that DGL is used for distributed graph neural network training, the graph structure data can be segmented according to the requirements of DGL for distributed graph neural network training to obtain multiple sub-graph data.

[0045] In an embodiment of the present application, sub-graph data obtained by segmenting graph structure data can be stored in a distributed memory management system, and data transmission between Spark and the graph neural network training architecture can be opened up through the distributed memory management system of shared memory.

[0046] In one example, the data after preprocessing the original data may also include feature data corresponding to the graph structure data. Accordingly, when segmenting the graph structure data, the feature data corresponding to the graph structure data may also be segmented to obtain feature data corresponding to each sub-graph data, and the sub-graph data and the feature data corresponding to the sub-graph data may be associated and stored.

[0047] For example, assuming that the graph structure data is segmented to obtain subgraph data 1 to 3 and the feature data corresponding to each subgraph data, then subgraph data 1 and its corresponding feature data, subgraph data 2 and its corresponding feature data, and subgraph data 3 and its corresponding feature data can be stored in the distributed memory management system. For example, subgraph data 1 and its corresponding feature data are stored in the memory of worker node 1 in the distributed memory management system, subgraph data 2 and its corresponding feature data are stored in the memory of worker node 2 in the distributed memory management system, and subgraph data 3 and its corresponding feature data are stored in the memory of worker node 3 in the distributed memory management system.

[0048] Step S120, construct a distributed graph neural network training function, use the distributed graph neural network training function, perform distributed graph neural network model training based on the subgraph data stored in the distributed memory management system, and store the obtained Embedding in the distributed memory management system.

[0049] In the embodiments of the present application, in order to implement distributed graph neural network training in Spark, a distributed graph neural network training function can be constructed, so as to utilize the constructed distributed graph neural network training function in Spark to perform distributed graph neural network model training based on the sub-graph data stored in the distributed memory management subsystem, and store the obtained Embedding into the distributed memory management system.

[0050] Exemplarily, a distributed graph neural network training function can be constructed according to a predefined graph neural network model and training parameters associated with the graph neural network model.

[0051] In one example, the logic of the distributed graph neural network training function can include but is not limited to:

[0052] Initialization of distributed training, reading and construction of training data, sub-graph sampling for mini-batch model training and verification, printing of training logs, and storage of the inferred Embedding.

[0053] For example, the main logic of the distributed graph neural network training function can include:

[0054] 1) Introduce a distributed deep learning training framework, such as Horovod, for distributed training;

[0055] 2) Read the sub-graph data stored in the distributed memory management system and convert it into a Tensor (tensor) to construct a DGL Graph (graph) for graph neural network training;

[0056] 3) Construct the overall process of graph neural network mini-batch training, including: sub-graph sampling, forward calculation of loss, backward propagation calculation of gradients, model parameter update, saving of checkpoint (an internal event), and printing of training information, etc.;

[0057] 4) After training is completed, use the model to infer the Embedding and store it in the distributed memory management system.

[0058] Step S130: Train and predict the ML model based on the Embedding saved in the distributed memory management system.

[0059] In the embodiments of the present application, when the Embedding is obtained and saved to the distributed memory management system in the manner described in the above embodiments, Spark can be used to read the Embedding existing in the distributed memory management system and perform data format conversion, and use user-defined logic to complete the training and prediction of the ML model.

[0060] It can be seen that in Figure 1 In the method flow shown, by combining the big data processing framework and the graph neural network framework and introducing a shared memory distributed memory management system, the data transmission between Spark and the graph neural network training architecture is opened up, data fusion is efficiently realized, and frequent data movement between the big data processing framework and the graph neural network framework is avoided. The data interaction efficiency between the big data framework and the graph neural network framework is improved, thereby improving the execution efficiency of graph data mining.

[0061] In some embodiments, the same subgraph data is stored in the memory of the same worker node, and multiple subgraph data are stored in the memory of at least two different worker nodes.

[0062] Exemplarily, in order to improve the data acquisition efficiency of distributed graph neural network training, and thereby improve the efficiency of distributed graph neural network training, the same subgraph data is stored in the memory of the same worker node. The worker node can perform distributed graph neural network training based on the subgraph data stored in the memory of this node, without the need to obtain data from the memory of other worker nodes, thereby improving data acquisition efficiency.

[0063] Exemplarily, data of multiple subgraphs obtained by segmenting the graph structure data can be stored in the memory of at least two different nodes, so that distributed graph neural network training can be performed through the at least two different worker nodes to improve the efficiency of graph neural network training.

[0064] In one example, data of different subgraphs are stored in the memory of different worker nodes.

[0065] Exemplarily, in order to improve the efficiency of distributed graph neural network training and make full use of device performance, the multiple sub-graph data obtained by segmenting the graph structure data can be stored in the memory of different worker nodes respectively, and the worker node storing the sub-graph data can be used to perform distributed graph neural network training based on the sub-graph data stored in the memory of this node.

[0066] In another example, at least two sub-graph data are stored in the memory of at least one worker node, and the worker node runs at least two distributed graph neural network training processes, and one distributed graph neural network training process corresponds to one sub-graph data.

[0067] Exemplarily, in order to make full use of device resources, when there is sufficient resources in a single worker node, at least two sub-graph data can be stored in the memory of the single worker node. The worker node can run at least two distributed graph neural network training processes, with one distributed graph neural network training process corresponding to one sub-graph data, and distributed graph neural network training is performed based on the sub-graph data.

[0068] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of the present application, the technical solutions provided in the embodiments of the present application will be described below with specific examples.

[0069] As Figure 2 shown, it is a schematic architecture diagram of a data processing system implemented based on the Apache Spark architecture provided in the embodiments of the present application. As Figure 2 shown, the data processing system may include a Driver (driver), a Cluster Manager (cluster manager), worker nodes, and a distributed memory management system.

[0070] In this embodiment, the big data processing framework Spark and the graph neural network framework DGL are combined, and a distributed memory management system with shared memory is introduced, so that preprocessing of graph data, training of a graph neural network model, obtaining Embedding, and training and prediction of Spark ML can be completed in a set of Spark Application.

[0071] The following is an explanation of the specific implementation process.

[0072] Please refer to Figure 3 which is a schematic flow diagram of a graph data mining method provided in the embodiments of the present application. As Figure 3 shown, the graph data mining method may include the following steps:

[0073] 1. Read the original data for preprocessing to obtain structured graph construction data (i.e., the above-mentioned graph structure data) and feature data.

[0074] Exemplarily, in a Spark program, a graph neural network model, a Spark ML model, and related training parameters can be defined, and relevant data can be read and preprocessed into the DataFrame or RDD (Resilient Distributed Dataset) format.

[0075] Exemplarily, the preprocessed data may include structured graph data (usually in the form of a binary tuple or a ternary tuple, where each tuple represents an edge on the graph and can be represented as two columns <source, destination> or three columns <source, relation, destination> in a DataFrame).

[0076] Optionally, the preprocessed data may further include feature data

[0077] Exemplarily, the feature data may include the attribute data of each node or relationship in the graph. For example, assuming the graph is a social network graph, each node can be a person, and its feature data may include the age, gender, etc. of the person.

[0078] Exemplarily, finally, by calling the fit_gnn() interface and passing in the graph construction data and feature data, the subsequent system will automatically complete the subgraph splitting, data format conversion and storage, and the training of the graph neural network.

[0079] 2. According to the requirements of DGL for distributed GNN training, split the graph construction data and feature data into subgraphs, and finally obtain the subgraph data required for distributed GNN training and the corresponding feature data of the subgraphs, etc.

[0080] Exemplarily, the splitting of DGL subgraphs first requires the allocation of partitions for each node according to a certain condition, and then includes the data required for each subgraph.

[0081] Exemplarily, when splitting the subgraph data, it is necessary to proceed based on the information of the entire graph, and the subgraph splitting should try to ensure that the relationships between nodes in the original graph are not damaged.

[0082] Exemplarily, the data may include but is not limited to: the original ID of the node, the new ID after reallocation, the local node flag bit, the local node features, labels, masks, and other information, as well as all the incident edges of the node.

[0083] For example, taking Figure 4 the graph structure shown as an example, assuming the entire graph is split into 3 subgraphs, and the splitting strategy is random splitting, that is, using the random allocation of nodes, nodes 1, 4, and 7 are allocated to the first subgraph; nodes 2, 3, and 5 are allocated to the second subgraph; nodes 6, 8, and 9 are allocated to the third subgraph. Then the structures of each subgraph are as shown in the figure: taking the first subgraph as an example, the other uncolored nodes are maintained by other subgraphs, and the nodes 1, 4, and 7 marked in blue are all the nodes (local nodes) belonging to this subgraph. It is necessary to extract the feature, label, mask, and other data of these local nodes and store them together with the topological relationship of the first subgraph in the distributed memory management system.

[0084] In addition, each node in Sub - graph 1 needs to maintain its new ID after global allocation. For example, the new ID of Node 1 is 1; the new ID of Node 4 is 2; the new ID of Node 7 is 3; the new ID of Node 2 is 4; the new ID of Node 3 is 5; the new ID of Node 6 is 7; the new ID of Node 8 is 8

[0085] The same applies to Sub - graph 2 and Sub - graph 3.

[0086] It should be noted that the allocation of new IDs can be based on the division of sub - graphs and can be globally allocated according to the sub - graphs to which the nodes belong.

[0087] Take Figure 4 the shown splitting result as an example. According to the sub - graphs to which each node belongs, the nodes can be sorted by ID respectively: that is, 1, 4, 7 in Sub - graph 1, 2, 3, 5 in Sub - graph 2, and 6, 8, 9 in Sub - graph 3. Then, new IDs can be allocated to each node in turn. That is, the new IDs of 1, 4, 7, 2, 3, 5, 6, 8, 9 are 1, 2, 3, 4, 5, 6, 7, 8, 9 in sequence.

[0088] The nodes split into each sub - graph

[0089] Exemplarily, use Spark to load the whole graph, and then the above - mentioned sub - graphs can be obtained according to a series of operators provided by Spark. The final splitting effect is to maintain the sub - graph data corresponding to each part in the Executor (executor) of each worker node.

[0090] 3. Convert and store the data formats of the sub - graph data and the feature data corresponding to the sub - graph.

[0091] Exemplarily, for the pre - processed data (i.e., the sub - graph data and the feature data corresponding to the sub - graph) in each data partition, convert it into the Apache Arrow format and store it in the distributed memory management system, and ensure that the aforementioned sub - graph data and the corresponding feature data are stored in the corresponding nodes of the distributed memory management system to achieve memory sharing.

[0092] 4. Construct a function for distributed GNN model training (i.e., the above - mentioned distributed graph neural network training function) according to the user - defined GNN model and related training parameters.

[0093] Exemplarily, the main logic of the distributed GNN model training function can include:

[0094] Initialization of distributed training, reading and construction of training data, mini - batch model training and verification by sub - graph sampling, printing of training logs, and storage of Embedding obtained by inference, etc.

[0095] Exemplarily, the distributed GNN model training function can be called through a Python process in a Spark Task.

[0096] 5. Use Horovod to start distributed DGL in a Spark Task. Each training process reads the subgraph data stored in the distributed memory management system described above, completes the construction of the graph and the Embedding training and inference of the graph neural network model, and stores the obtained Embedding in the distributed memory management system.

[0097] Exemplarily, an RDD can be constructed according to the number of GPU cards set for DGL distributed training. The number of partitions of the RDD is set to be the same as the number of GPU cards required for GNN distributed training, so as to ensure that one Task is bound to one GPU. Then, start a Task on the RDD through the mapPartition (mapping partition, a kind of mapping function) method. What the Task actually executes is the Horovod distributed training, that is, the function for training the graph neural network constructed in step 4 (i.e., the above-mentioned distributed GNN model training function).

[0098] 6. Read the Embedding stored in the distributed memory management system, and use the user-defined Spark ML model to complete model training and achieve classification prediction.

[0099] Exemplarily, after the Embedding is obtained through distributed training and written into the memory management system, read the Embedding in the distributed memory management system in Spark and perform data format conversion, and use the user-defined logic to complete the training and prediction of the Spark ML model.

[0100] The method provided by the present application has been described above. Next, the device provided by the present application will be described:

[0101] Please refer to Figure 5 , which is a schematic structural diagram of a graph data mining device provided by an embodiment of the present application. As Figure 5 shown, the graph data mining device may include:

[0102] A preprocessing unit 510, configured to preprocess the original data to obtain graph structure data;

[0103] A splitting unit 520, configured to split the graph structure data according to the training strategy of the distributed graph neural network to obtain a plurality of subgraph data, and store the subgraph data in the distributed memory management system;

[0104] A training unit 530 for constructing a distributed graph neural network training function, using the distributed graph neural network training function to perform distributed graph neural network model training based on the subgraph data stored in the distributed memory management system, and storing the obtained Embedding in the distributed memory management system;

[0105] A mining unit 540 for training and predicting an ML model based on the Embedding stored in the distributed memory management system.

[0106] In some embodiments, the same subgraph data is stored in the memory of the same working worker node, and the multiple subgraph data is stored in the memories of at least two different worker nodes.

[0107] In some embodiments, the data of different subgraphs is stored in the memories of different worker nodes;

[0108] Or,

[0109] At least two subgraph data are stored in the memory of at least one worker node, and at least two distributed graph neural network training processes are running on this worker node, with one distributed graph neural network training process corresponding to one subgraph data.

[0110] In some embodiments, the splitting unit 520 splits the graph structure data according to the training strategy of the distributed graph neural network, including:

[0111] Splitting the graph structure data according to the requirements of the target graph deep learning framework for distributed graph neural network training; wherein, the target graph deep learning framework is a graph deep learning framework pre-configured for distributed graph neural network training.

[0112] In some embodiments, the training unit 530 constructs a distributed graph neural network training function, including:

[0113] Constructing the distributed graph neural network training function according to a predefined graph neural network model and training parameters associated with the graph neural network model;

[0114] Wherein, the logic of the distributed graph neural network training function includes:

[0115] Initialization of distributed training, reading and construction of training data, mini-batch model training and verification by subgraph sampling, printing of training logs, and storage of the obtained Embedding by inference.

[0116] In some embodiments, the preprocessing unit 510 preprocesses the original data to obtain graph structure data, including:

[0117] Preprocess the original data to obtain graph structure data and feature data corresponding to the graph structure data;

[0118] The splitting unit 520 splits the graph structure data according to the training strategy of the distributed graph neural network to obtain a plurality of sub-graph data, and stores the sub-graph data in a distributed memory management system, including:

[0119] Split the graph structure data and the feature data corresponding to the graph structure data according to the training strategy of the distributed graph neural network to obtain a plurality of sub-graph data and the feature data corresponding to each sub-graph data, and store the sub-graph data and the feature data corresponding to the sub-graph data in an associated manner.

[0120] Please refer to Figure 6 , which is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. The electronic device may include a processor 601 and a memory 602 storing machine-executable instructions. The processor 601 and the memory 602 may communicate via a system bus 603. And by reading and executing the machine-executable instructions corresponding to the graph data mining control logic in the memory 602, the processor 601 may execute the graph data mining method described above.

[0121] The memory 602 mentioned herein may be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, and so on. For example, the machine-readable storage medium may be: RAM (Radom Access Memory, random access memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.

[0122] In some embodiments, a machine-readable storage medium is also provided, such as Figure 6 the memory 602 in, which stores machine-executable instructions, and when the machine-executable instructions are executed by a processor, the graph data mining method described above is implemented. For example, the machine-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc.

[0123] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0124] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A graph data mining method, characterized in that, Applied to a data processing system implemented based on the Apache Spark architecture, the method includes: Preprocess the original data to obtain graph-structured data; According to the training strategy of the distributed graph neural network, split the graph-structured data to obtain multiple sub-graph data, and store the sub-graph data in the distributed memory management system; Construct a distributed graph neural network training function, and use the distributed graph neural network training function to perform distributed graph neural network model training based on the sub-graph data stored in the distributed memory management system, and store the obtained representation Embedding in the distributed memory management system; Train and predict the machine learning ML model based on the Embedding saved in the distributed memory management system; Wherein, the preprocessing the original data to obtain graph-structured data includes: Preprocess the original data to obtain graph-structured data and the feature data corresponding to the graph-structured data; The splitting the graph-structured data according to the training strategy of the distributed graph neural network to obtain multiple sub-graph data and storing the sub-graph data in the distributed memory management system includes: According to the training strategy of the distributed graph neural network, split the graph-structured data and the feature data corresponding to the graph-structured data to obtain multiple sub-graph data and the feature data corresponding to each sub-graph data, and store the sub-graph data and the feature data corresponding to the sub-graph data in an associated manner.

2. The method according to claim 1, wherein The same sub-graph data is stored in the memory of the same working worker node, and the multiple sub-graph data are stored in the memories of at least two different worker nodes.

3. The method according to claim 2, wherein The data of different sub-graphs are stored in the memories of different worker nodes; Or, At least two sub-graph data are stored in the memory of at least one worker node, and at least two distributed graph neural network training processes are running on this worker node, and one distributed graph neural network training process corresponds to one sub-graph data.

4. The method according to claim 1, characterized in that, The splitting the graph-structured data according to the training strategy of the distributed graph neural network includes: Split the graph-structured data according to the requirements of the target graph deep learning framework for distributed graph neural network training; wherein, the target graph deep learning framework is a graph deep learning framework pre-configured for distributed graph neural network training.

5. The method according to claim 1, characterized in that, The constructing the distributed graph neural network training function includes: Construct the distributed graph neural network training function according to the predefined graph neural network model and the training parameters associated with the graph neural network model; Wherein, the logic of the distributed graph neural network training function includes: Initialization of distributed training, reading and construction of training data, mini-batch model training and verification by sub-graph sampling, printing of training logs, and storage of the obtained Embedding for inference.

6. A graph data mining device, characterized in that, Applied to a data processing system implemented based on the Apache Spark architecture, the device includes: A preprocessing unit for preprocessing the original data to obtain graph-structured data; A splitting unit, configured to split the graph structure data according to the training strategy of the distributed graph neural network, obtain multiple sub-graph data, and store the sub-graph data in a distributed memory management system; A training unit, configured to construct a distributed graph neural network training function, and use the distributed graph neural network training function to perform distributed graph neural network model training according to the sub-graph data stored in the distributed memory management system, and store the obtained representation Embedding in the distributed memory management system; A mining unit, configured to perform training and prediction of a machine learning ML model according to the Embedding stored in the distributed memory management system; Wherein, the preprocessing unit preprocesses the original data to obtain graph structure data, including: Preprocessing the original data to obtain graph structure data and feature data corresponding to the graph structure data; The splitting unit splits the graph structure data according to the training strategy of the distributed graph neural network, obtains multiple sub-graph data, and stores the sub-graph data in the distributed memory management system, including: Splitting the graph structure data and the feature data corresponding to the graph structure data according to the training strategy of the distributed graph neural network, obtaining multiple sub-graph data and feature data corresponding to each sub-graph data, and associatively storing the sub-graph data and the feature data corresponding to the sub-graph data.

7. The device according to claim 6, characterized in that, The same sub-graph data is stored in the memory of the same working worker node, and the multiple sub-graph data are stored in the memories of at least two different worker nodes; Wherein, the data of different sub-graphs are stored in the memories of different worker nodes; Or, At least two sub-graph data are stored in the memory of at least one worker node, and at least two distributed graph neural network training processes are running on this worker node, and one distributed graph neural network training process corresponds to one sub-graph data; And / or, the splitting unit splits the graph structure data according to the training strategy of the distributed graph neural network, including: Splitting the graph structure data according to the requirements of the target graph deep learning framework for distributed graph neural network training; wherein, the target graph deep learning framework is a graph deep learning framework configured in advance for distributed graph neural network training; And / or, the training unit constructs a distributed graph neural network training function, including: Constructing the distributed graph neural network training function according to a predefined graph neural network model and training parameters associated with the graph neural network model; Wherein, the logic of the distributed graph neural network training function includes: Initialization of distributed training, reading and construction of training data, mini-batch model training and verification by sub-graph sampling, printing of training logs, and storage of the obtained Embedding by inference.

8. An electronic device, characterized in that, It includes a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor is used to execute the machine-executable instructions to implement the method according to any one of claims 1-5.

9. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by a processor, the method according to any one of claims 1-5 is implemented.