A Big Data Relationship Mining Method Based on Graph Computing

By using graph computing technology, deep learning models and graph convolutional networks in big data relationship mining, the performance bottlenecks of large-scale graph data processing in the existing technology and the insufficient depth and breadth of relationship mining are solved, and more efficient sub-graph mining and more accurate relationship extraction are achieved, which improves the effect of data analysis.

CN119312810BActive Publication Date: 2025-05-30GUANGZHOU JINGCUI EDUCATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411304787.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2025-05-30
Estimated Expiration
2044-09-19

AI Technical Summary

Technical Problem

The existing technology faces computing performance bottlenecks in large-scale graph data processing, which cannot effectively improve the efficiency of sub-graph mining and pattern matching, and the depth and breadth are poor when mining character relationships, and the management and update effect is poor.

Method used

The graph-based calculation method is adopted to improve the accuracy of relationship extraction through deep learning models (such as entity recognition model combined with Bi-LSTM and CRF and BERT relationship extraction model), and to process large-scale graph data using graph convolutional network (GCN) and distributed graph computing framework (such as GraphX) to improve the efficiency of subgraph mining and pattern matching. At the same time, data deduplication and noise filtering technologies are used to improve data quality.

Benefits of technology

It improves the accuracy of relationship extraction and the efficiency of sub-graph mining, improves data quality, provides more reliable input for subsequent relationship mining and pattern analysis, and enhances the intuitiveness and practicality of data analysis through the visualization technology of interactive graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119312810B_ABST
    Figure CN119312810B_ABST
Patent Text Reader

Abstract

The present invention proposes a big data relationship mining method based on graph computing. By preprocessing multi-source heterogeneous data to generate a graph structure, it uses graph computing technology to mine frequent subgraphs and dynamic relationships, and constructs a relationship model. The method involves data cleaning, entity recognition, relationship extraction, and graph computing technology. Among them, data cleaning combines locality-sensitive hashing and machine learning denoising, entity recognition and relationship extraction adopt deep learning models, and graph computing integrates graph convolutional networks and distributed graph traversal. Finally, relationship prediction is carried out through a graph-based machine learning framework, and the analysis results are displayed using graph visualization technology to provide in-depth data insights. This method effectively improves the accuracy and efficiency of big data relationship mining and is applicable to data analysis and decision support in complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing and mining technologies, and more specifically, to a big data relationship mining method based on graph computing. Background Art

[0002] In the big data era, the types and quantities of data have grown exponentially. Traditional data processing and mining methods have difficulty dealing with these complex data relationships and massive data. As an effective tool for processing complex data relationships and structures, graph computing has gradually attracted wide attention from researchers and the industry. Graph computing is a computing model based on graph theory for processing graph-structured data composed of nodes and edges. It is applicable to various application scenarios, including social network analysis, recommendation systems, financial risk assessment, etc. Graph computing can effectively capture the relationships and dependencies in data and reveal potential patterns and trends.

[0003] The prior art discloses a visualization person relationship mining management system and method based on big data. The visualization person relationship mining management system based on big data belongs to the field of big data technology and includes a mining module, an analysis module, and a management module. The mining module includes a person mining unit, a time mining unit, and a frequency mining unit. The person mining unit is used to count and calculate different accounts in the obtained mining dataset to obtain a person dataset. The time mining unit is used to count and mark the conversation time points in the person dataset to obtain a time point set. The duration of the conversation is obtained according to the conversation time point and counted and marked to obtain a duration set. The time point set and the duration set constitute a time dataset. The present invention also discloses a visualization person relationship mining management method based on big data. The present invention can solve the technical problems of poor depth and breadth in person relationship mining and poor management and update effects after person relationship mining in the existing solutions.

[0004] However, this data mining method may not be able to fully utilize the advantages of deep learning models during relationship extraction. When dealing with large-scale graph data, it faces computational performance bottlenecks and cannot improve the efficiency of subgraph mining and pattern matching. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, the present invention designs a big data relationship mining method based on graph computing, which can effectively overcome the defects of the prior art.

[0006] To solve the above technical problems, the technical solution of the present invention is as follows:

[0007] A big data relationship mining method based on graph computing includes the following steps:

[0008] S1. Preprocess the multi-source heterogeneous big data, including data cleaning, entity recognition, relationship extraction, and type annotation, to generate graph-structured data. The preprocessing step includes using natural language processing techniques to perform semantic role annotation on text data to identify the semantic relationships between entities;

[0009] S2. Use graph computing techniques to mine frequent subgraphs and dynamic relationships from the graph. The graph computing techniques include a pattern matching algorithm that combines deep learning feature extraction and traditional graph traversal techniques to improve the accuracy and efficiency of subgraph mining;

[0010] S3. Build a relationship model based on the mining results and apply it to specific scenarios for data analysis and insight. The relationship model includes a graph-based machine learning framework for training and predicting the development and changes of complex relationships between entities.

[0011] The data cleaning step includes using data deduplication and noise filtering techniques. The data deduplication uses the locality-sensitive hashing method, which includes the following steps:

[0012] Convert each data record into a feature vector;

[0013] Select a hash function suitable for the data characteristics;

[0014] Use the selected hash function to map each feature vector to one or more hash buckets. Each hash bucket contains multiple feature vectors, and the feature vectors in each hash bucket are relatively close in the hash space;

[0015] For each new data record, calculate the hash value of its feature vector and map it to the corresponding hash bucket;

[0016] Perform duplicate detection in the hash bucket, compare the similarity of all feature vectors in the same hash bucket, identify the records with similarity higher than the set threshold as duplicate data, and remove the identified duplicate data from the dataset, retaining only unique data records;

[0017] The noise filtering technique includes using statistical-based methods and machine learning-based methods, specifically including the following steps:

[0018] The statistical method includes: standard deviation analysis, calculating the standard deviation of each data, and identifying those data points with a large deviation from the mean;

[0019] Distribution anomaly detection, using a box plot to identify outliers in the data;

[0020] Applying the Z-Score method, data points with a Z-Score greater than 3 or less than -3 are considered outliers;

[0021] The machine learning-based method uses the Local Outlier Factor, specifically:

[0022] Distance calculation: Calculate the neighborhood distance using the Euclidean distance. For each data point, calculate its distance to all other data points;

[0023] Construct local density: Select the k value to determine the size of the local neighborhood (k value), that is, the k nearest neighbors of each data point;

[0024] Calculate the k-distance: For each data point, find the maximum distance (k-distance) among its k nearest neighbors;

[0025] Calculate the reachability distance: For data points i and j, calculate their reachability distance, defined as the larger value between the k-distances of i and j;

[0026] Local density: Calculate the local density of each data point, that is, the density of data points within its k neighborhoods;

[0027] Relative density: Calculate the relative density of each data point, that is, the ratio of its local density to the local density of data points within its neighborhood;

[0028] LOF value: Calculate the LOF value of each data point, that is, the average value of its relative density. Set a threshold. According to the LOF value, set a threshold. Data points with an LOF value exceeding the threshold are considered outliers;

[0029] Filter outliers: Remove or mark these outliers from the dataset.

[0030] The entity recognition step includes using a deep learning-based named entity recognition (NER) model, where the recognition model is a model combining a bidirectional long short-term memory network (Bi-LSTM) and a conditional random field (CRF). The specific method steps are as follows:

[0031] Obtain a labeled text dataset that contains labeled entity information;

[0032] Construct a vocabulary and a label set. The vocabulary maps the words in the text to unique index values;

[0033] The label set defines entity label categories;

[0034] Use pre-trained word embeddings (such as Word2Vec, GloVe) or a custom word embedding layer to convert words into vectors;

[0035] Bi-LSTM network:

[0036] Input layer, which takes word embeddings as the input of the Bi-LSTM. Bi-LSTM layer, which uses bidirectional LSTM to capture context information;

[0037] CRF layer:

[0038] Sequence labeling, which uses the CRF layer to perform sequence labeling on the output of the Bi-LSTM to optimize entity boundary recognition;

[0039] Model training, which uses the log-likelihood of the CRF as the loss function and the Adam optimization algorithm to train the model. The data is divided into a training set and a validation set for model training, and hyperparameters are adjusted to optimize performance.

[0040] The relationship extraction step includes using a deep learning-based relationship extraction model, which uses a pre-trained language model (such as BERT) for relationship prediction, and includes the following steps:

[0041] Obtain a large-scale corpus containing entity pairs and relationship annotations;

[0042] Organize the text into a format suitable for model input, and divide the dataset into a training set, a validation set, and a test set;

[0043] Text tokenization, which uses the tokenizer of BERT to split the text into vocabulary indices;

[0044] Generate input, create input samples for each entity pair, including:

[0045] Token IDs, the vocabulary indices of the text;

[0046] Attention Masks, indicating which positions should be attended to by the model;

[0047] Token Type IDs, used to distinguish two different text parts;

[0048] Model construction, load the pre-trained BERT and select a suitable BERT variant;

[0049] BERT layer, which extracts context information from the text;

[0050] Relationship classification head, add a fully connected layer on the CLS token output by BERT for relationship classification;

[0051] Define the loss function, use the cross-entropy loss function to optimize the classification task;

[0052] Perform model training and input new text into the trained model for relationship prediction, and map the relationship categories output by the model back to the actual relationship labels.

[0053] The pattern matching algorithm in the graph computing technology includes using a graph convolutional network (GCN) to perform feature learning and pattern recognition on graph-structured data, where the GCN is used to learn local and global features of nodes. The specific steps are as follows:

[0054] Graph data collection, obtaining graph-structured data with nodes and edges;

[0055] Data formatting, converting the data into a format suitable for GCN input, including a node feature matrix and an adjacency matrix, and normalizing the node features to improve the training effect;

[0056] Adjacency matrix processing, normalizing the adjacency matrix and calculating the inverse square root of the degree matrix of the adjacency matrix of each node;

[0057] Convolution operation, applying a convolution operation to each node and its neighborhood to aggregate the information of neighboring nodes;

[0058] Activation function, using the ReLU activation function to increase the non-linear expression ability of the model;

[0059] Stacking GCN layers, stacking multiple graph convolutional layers as needed to capture multi-level features of nodes;

[0060] Output layer, adding an appropriate output layer for the subgraph mining task, such as a fully connected layer or a pooling layer, and training the constructed model;

[0061] Feature extraction, extracting the features of nodes and subgraphs from the GCN output;

[0062] Subgraph matching, identifying and extracting subgraphs that conform to a specific pattern by comparing the extracted features with the target pattern;

[0063] Deploy the trained GCN model to the production environment, integrate it into the actual application, use the model for real-time subgraph matching and pattern recognition, and provide data analysis and decision support.

[0064] The graph traversal technology includes using a distributed graph computing framework (GraphX) to process large-scale graph data to accelerate the subgraph mining process;

[0065] Confirm the Spark cluster version and download the GraphX package compatible with the Spark version;

[0066] Upload the GraphX package to all nodes of the Spark cluster and place it in the specified library directory;

[0067] Adjust the spark.executor.memory and spark.driver.memory parameters to allocate appropriate memory;

[0068] Set the spark.cores.max and spark.executor.instances parameters to configure the number of cores and executor instances;

[0069] Data loading, use Spark's textFile or sequenceFile methods to load data from HDFS, S3 or other storage systems;

[0070] Data cleaning, remove invalid or incorrect records, and convert the original data into vertex format and edge format;

[0071] Create vertex RDD and edge RDD, ensuring that the data types are consistent with the requirements of GraphX;

[0072] Create vertex RDD, use Spark's parallelize method to create vertex RDD, and assign a unique ID and corresponding attributes to each vertex;

[0073] Create edge RDD, use Spark's parallelize method to create edge RDD, and define the source vertex, target vertex and attributes of the edge;

[0074] Build a Graph object, pass the vertex RDD and edge RDD to the Graph constructor to create a Graph object, and set the default vertex attributes for the Graph;

[0075] Define the graph calculation logic:

[0076] Select a suitable graph algorithm according to business requirements and implement the algorithm logic;

[0077] Use the GraphX API, call the pregel and mapVertices APIs of GraphX for graph calculation;

[0078] Write user-defined functions to process vertex and edge attributes;

[0079] Distributed graph calculation

[0080] Graph partitioning, select a suitable partitioning strategy;

[0081] Consider the characteristics of the graph to select the best partitioning strategy;

[0082] Parallel computing, utilize the distributed computing power of GraphX, split the graph data into multiple partitions and process them in parallel on the cluster, monitor the progress of tasks and resource usage to ensure the efficient operation of computing tasks.

[0083] The graph-based machine learning framework in the relational model includes using a graph neural network (GNN) to model and predict complex relationships between entities, where the GNN is used to capture the dependencies and dynamic changes between entities.

[0084] The data analysis and insight step includes using graph-based visualization techniques to present the mined relationships and patterns as interactive graphs for users to conduct in-depth analysis and understanding.

[0085] A computer storage medium stores computer instructions that, when called, are used to execute the above-mentioned big data relationship mining method based on graph computing.

[0086] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows: A big data relationship mining method based on graph computing uses deep learning models (such as the combination of Bi-LSTM and CRF for entity recognition, and BERT for relation extraction) to improve the model's ability to understand complex languages and relationships, thereby improving the accuracy of relation extraction. This approach makes up for the deficiencies of the prior art methods in semantic understanding, enabling the model to better capture the deep relationships in the text. The graph convolutional network (GCN) and distributed graph computing framework (such as GraphX) are used to process large-scale graph data, which solves the performance bottleneck of traditional methods when dealing with big data. The use of distributed computing improves the efficiency of subgraph mining and pattern matching. Through data deduplication and noise filtering techniques (including locality-sensitive hashing, statistical methods, and machine learning methods), the quality of the data is effectively improved, providing a more reliable input for subsequent mining. The graph-based visualization technique presents the mined relationships and patterns as interactive graphs, helping users to more deeply understand the complex relationships and patterns in the data, and enhancing the intuitiveness and practicality of data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained based on the provided drawings.

[0088] Figure 1 It is a flowchart of the steps of a big data relationship mining method based on graph computing. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0089] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0090] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, which do not represent the dimensions of the actual product;

[0091] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0092] The technical solution of the present invention will be further described below in conjunction with the drawings and embodiments.

[0093] Embodiment

[0094] A big data relationship mining method based on graph computing, as Figure 1 shown, includes the following steps:

[0095] S1. Preprocess multi-source heterogeneous big data, including data cleaning, entity recognition, relationship extraction, and type annotation, to generate graph-structured data, where the preprocessing step includes using natural language processing technology to perform semantic role annotation on text data to identify the semantic relationships between entities;

[0096] S2. Use graph computing technology to mine frequent subgraphs and dynamic relationships from the graph, where the graph computing technology includes a pattern matching algorithm, and the pattern matching algorithm combines the feature extraction of deep learning and traditional graph traversal technology to improve the accuracy and efficiency of subgraph mining;

[0097] S3. Build a relationship model based on the mining results and apply it to specific scenarios for data analysis and insight, where the relationship model includes a graph-based machine learning framework for training and predicting the development and changes of complex relationships between entities.

[0098] The data cleaning step includes using data deduplication and noise filtering technologies. The data deduplication uses the locality-sensitive hashing method, which includes the following steps:

[0099] Convert each data record into a feature vector;

[0100] Select a hash function suitable for the data characteristics; such as the MinHash hash function for Euclidean distance or the LSH hash function for cosine similarity, such as LSH for cosine similarity.

[0101] Use the selected hash function to map each of the feature vectors into one or more hash buckets, each of the hash buckets contains multiple of the feature vectors, and the feature vectors of each of the hash buckets are relatively close in the hash space;

[0102] For each new data record, calculate the hash value of its feature vector and map it to the corresponding hash bucket;

[0103] Perform duplicate detection in the hash bucket, compare the similarity of all the feature vectors in the same hash bucket, identify the records with similarity higher than the set threshold as duplicate data, remove the identified duplicate data from the dataset, and retain the unique data records;

[0104] The noise filtering technology includes using statistical-based methods and machine learning-based methods, and specifically includes the following steps:

[0105] The statistical methods include: standard deviation analysis, calculating the standard deviation of each data, and identifying the data points with large deviations from the mean;

[0106] Distribution anomaly detection, using a box plot (Box-plot) to identify outliers in the data;

[0107] A box plot (Box-plot), also known as a box-and-whisker plot, is a statistical chart used to display the distribution of a set of data. It can provide summary statistical information of the minimum value, first quartile (Q1), median (Q2), third quartile (Q3), and maximum value of the data, and can intuitively show the central tendency, dispersion degree, and outliers of the data through the shape of the box and the length of the whiskers.

[0108] Apply the Z-Score method, where data points with Z-Score greater than 3 or less than -3 are considered outliers;

[0109] The machine learning-based method uses the local outlier factor, specifically:

[0110] Distance calculation, using the Euclidean distance to calculate the neighborhood distance, and for each data point, calculate its distance from all other data points;

[0111] Construct local density, select the k value to determine the size of the local neighborhood (k value), that is, the k nearest neighbors of each data point;

[0112] Calculate the k-distance, and for each data point, find the maximum distance (k-distance) among its k nearest neighbors;

[0113] Calculate the reachable distance, and for data points i and j, calculate their reachable distance, which is defined as the larger value between the k-distance of i and j;

[0114] Local density, calculate the local density of each data point, that is, the density of the data points within its k neighborhoods;

[0115] Relative density, calculate the relative density of each data point, that is, the ratio of its local density to the local density of the data points within its neighborhood;

[0116] Calculate the LOF value for each data point, which is the average of its relative density. The larger the LOF value, the more likely the data point is an outlier. Set a threshold based on the LOF value, and data points with LOF values exceeding the threshold are considered outliers.

[0117] Filter out the outliers by removing or marking these outliers from the dataset.

[0118] The entity recognition step includes using a deep learning-based named entity recognition (NER) model, where the recognition model is a model that combines a bidirectional long short-term memory network (Bi-LSTM) and a conditional random field (CRF). The specific method steps are as follows:

[0119] Obtain a labeled text dataset that contains labeled entity information, such as person names, place names, and organization names.

[0120] Construct a vocabulary and a label set. The vocabulary maps the words in the text to unique index values.

[0121] The label set defines entity label categories. For example, "B-PER" represents the beginning of a person name, and "I-PER" represents the inside of a person name.

[0122] Convert words into vectors using pre-trained word embeddings (such as Word2Vec, GloVe) or a custom word embedding layer.

[0123] Bi-LSTM network:

[0124] Input layer: Take the word embedding as the input of the Bi-LSTM. Bi-LSTM layer: Use bidirectional LSTM to capture context information.

[0125] CRF layer:

[0126] Sequence labeling: Use the CRF layer to perform sequence labeling on the output of the Bi-LSTM to optimize entity boundary recognition.

[0127] Model training: Use the log-likelihood of the CRF as the loss function and use the Adam optimization algorithm to train the model. Divide the data into a training set and a validation set, perform model training, and adjust the hyperparameters to optimize the performance.

[0128] The relation extraction step includes using a deep learning-based relation extraction model that uses a pre-trained language model (such as BERT) for relation prediction. It includes the following steps:

[0129] Obtain a large-scale corpus that contains entity pairs and relation annotations, such as the ACE and SemEval datasets.

[0130] Organize the text into a format suitable for model input, and divide the dataset into training set, validation set, and test set;

[0131] Text tokenization, using the tokenizer of BERT to split the text into vocabulary indices;

[0132] Generate inputs, create input samples for each entity pair, including:

[0133] Token IDs, the vocabulary indices of the text;

[0134] Attention Masks, indicating which positions should be attended to by the model;

[0135] Token Type IDs, used to distinguish two different text parts; such as questions and answers.

[0136] Model construction, load the pre-trained BERT, and select a suitable BERT variant, such as BERT-base or BERT-large.

[0137] BERT layer, extract the context information in the text;

[0138] Relationship classification head, add a fully connected layer on the CLS token output by BERT for relationship classification;

[0139] Define the loss function, use the cross-entropy loss function to optimize the classification task;

[0140] Train the model and input new text into the trained model for relationship prediction, and map the relationship categories output by the model back to the actual relationship labels.

[0141] The pattern matching algorithm in the graph computing technology includes using a graph convolutional network (GCN) to perform feature learning and pattern recognition on graph-structured data, where the GCN is used to learn the local and global features of nodes. The specific steps are as follows:

[0142] Graph data collection, obtain graph-structured data with nodes and edges; (The data can be sourced from social networks, molecular structures, or other graph-formatted datasets

[0143] Data formatting, convert the data into a format suitable for GCN input, including node feature matrices and adjacency matrices, and standardize the node features to improve the training effect;

[0144] Adjacency matrix processing, perform normalization on the adjacency matrix, and calculate the inverse square root of the degree matrix of the adjacency matrix of each node;

[0145] Convolution operation, apply convolution operations to each node and its neighborhood to aggregate the information of neighboring nodes;

[0146] Activation function, use the ReLU activation function to increase the non - linear expression ability of the model;

[0147] Stack GCN layers, stack multiple graph convolutional layers as needed to capture multi - level features of nodes;

[0148] Output layer, add an appropriate output layer for the sub - graph mining task, such as a fully - connected layer or a pooling layer, and train the constructed model;

[0149] Feature extraction, extract features of nodes and sub - graphs from the GCN output;

[0150] Sub - graph matching, identify and extract sub - graphs that match a specific pattern by comparing the extracted features with the target pattern;

[0151] Deploy the trained GCN model to the production environment, integrate it into the actual application, use the model for real - time sub - graph matching and pattern recognition, and provide data analysis and decision - making support.

[0152] The graph traversal technology includes using a distributed graph computing framework (GraphX) to process large - scale graph data to accelerate the sub - graph mining process;

[0153] Confirm the Spark cluster version, download the GraphX package compatible with the Spark version;

[0154] Upload the GraphX package to all nodes of the Spark cluster and place it in the specified library directory;

[0155] Adjust the spark.executor.memory and spark.driver.memory parameters to allocate appropriate memory;

[0156] Set the spark.cores.max and spark.executor.instances parameters to configure the number of cores and executor instances;

[0157] Data loading, use Spark's textFile or sequenceFile method to load data from HDFS, S3 or other storage systems;

[0158] Data cleaning, remove invalid or incorrect records, and convert the original data into vertex format ((VertexId, VD)) and edge format (Edge[ED]);

[0159] Create vertex RDD and edge RDD, ensure that the data types are consistent with the requirements of GraphX;

[0160] Create a vertex RDD. Use Spark's parallelize method to create a vertex RDD, and assign a unique ID and corresponding attributes to each vertex;

[0161] Create an edge RDD. Use Spark's parallelize method to create an edge RDD, and define the source vertex, target vertex, and attributes of the edge;

[0162] Build a Graph object. Pass the vertex RDD and edge RDD to the Graph constructor to create a Graph object, and set the default vertex attributes for the Graph;

[0163] Define the graph computation logic:

[0164] Select a suitable graph algorithm according to business requirements and implement the algorithm logic; for example, define the iterative calculation process of PageRank;

[0165] Use the GraphX API. Call the pregel and mapVertices APIs of GraphX for graph computation;

[0166] Write user-defined functions to process the attributes of vertices and edges;

[0167] Distributed graph computation

[0168] Graph partitioning. Select a suitable partitioning strategy; such as CanonicalRandomVertexCut, to optimize the computation performance;

[0169] Consider the characteristics of the graph to select the best partitioning strategy;

[0170] For a connected graph, CanonicalRandomVertexCut usually performs well;

[0171] For a sparse graph or a dense graph, EdgeCut may be more suitable for a dense graph, while VertexCut is suitable for a sparse graph.

[0172] Parallel computation. Utilize the distributed computing ability of GraphX to split the graph data into multiple partitions and process them in parallel on the cluster, monitor the progress of tasks and resource usage, and ensure the efficient operation of the computation tasks.

[0173] The graph-based machine learning framework in the relational model includes using a graph neural network (GNN) to model and predict complex relationships between entities, where the GNN is used to capture the dependencies and dynamic changes between entities.

[0174] The data analysis and insight step includes using graph-based visualization techniques to present the mined relationships and patterns as interactive graphs for users to conduct in-depth analysis and understanding.

[0175] A computer storage medium stores computer instructions which, when called, are used to execute the above-mentioned big data relationship mining method based on graph computing.

[0176] In the specific implementation process, step S1: Data cleaning and graph structure generation

[0177] Convert each data record into a feature vector, use the Locality-Sensitive Hashing (LSH) method for deduplication, select a hash function suitable for the data features, map the feature vector into a hash bucket, for each new data record, calculate its hash value and map it into the hash bucket, perform duplicate detection in the hash bucket, identify and remove duplicate data by calculating similarity, and retain unique records.

[0178] Noise filtering:

[0179] Statistical methods, use standard deviation analysis, box plots, and the Z-Score method to identify and remove outliers, machine learning methods: calculate the Local Outlier Factor (LOF) for each data point, set a threshold, and remove outliers with LOF values exceeding the threshold.

[0180] Entity recognition and relationship extraction:

[0181] Entity recognition, use a deep learning model combining Bidirectional Long Short-Term Memory Network (Bi-LSTM) and Conditional Random Field (CRF) for named entity recognition, construct a vocabulary and label set, convert words into vectors using pre-trained or custom word embeddings, the Bi-LSTM layer captures context information, and the CRF layer performs sequence labeling to optimize entity boundary recognition.

[0182] Relationship extraction, use a deep learning model based on BERT for relationship prediction, organize the text into a format suitable for model input, use BERT to extract context information, and add a fully connected layer for relationship classification.

[0183] Generate graph structure data, convert the preprocessed data (including entities and relationships) into graph structure data, including nodes (entities) and edges (relationships).

[0184] Step S2: Graph computing and frequent subgraph mining

[0185] Graph computing technology, Graph Convolutional Network (GCN): Collect and format graph data, create a node feature matrix and an adjacency matrix, normalize the adjacency matrix, apply convolutional operations and ReLU activation functions in the GCN, stack multiple GCN layers to capture multi-level features of nodes, extract features of nodes and subgraphs from the GCN output, and perform pattern matching and subgraph recognition.

[0186] A pattern matching algorithm uses a graph convolutional network (GCN) for pattern recognition, identifies subgraphs that match the target pattern, and extracts frequent subgraphs.

[0187] Step S3: Relationship model construction and application

[0188] Construct a relationship model, use a graph neural network (GNN) to model and predict complex relationships between entities, and train the graph neural network to capture the dependencies and dynamic changes between entities.

[0189] Apply the model, apply the trained relationship model to a specific scenario for data analysis and insights.

[0190] Step S4: Data analysis and visualization

[0191] Data analysis, use graph-based visualization techniques to present the mined relationships and patterns as an interactive graph. Through the graph, users can conduct in-depth analysis to understand the complex relationships and patterns in the data.

[0192] Visualization techniques, select appropriate graph visualization tools and platforms (such as Gephi, Cytoscape)

[0193] Design an interactive interface that allows users to query and analyze the nodes and edges in the graph, and provide functions for filtering, zooming, and layout of the graph for users to conduct flexible exploration and analysis.

[0194] Identical or similar reference numerals correspond to identical or similar components;

[0195] The terms describing the positional relationships in the drawings are for illustrative purposes only and should not be construed as a limitation of this patent;

[0196] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A big data relationship mining method based on graph computing, characterized in that: The following steps are involved: S1. Preprocessing multi-source heterogeneous big data, including data cleaning, entity recognition, relationship extraction and type labeling, to generate graph structure data, wherein the preprocessing step includes using natural language processing technology to perform semantic role labeling on text data to identify semantic relationships between entities; S2. mining frequent subgraphs and dynamic relationships from a graph using graph computing technology, wherein the graph computing technology includes a pattern matching algorithm that combines feature extraction from deep learning with traditional graph traversal technology; The pattern matching algorithm in the graph computing technology includes using a graph convolutional network to perform feature learning and pattern recognition on graph structure data, wherein the graph convolutional network is used to learn local and global features of nodes, and the specific steps are: Graph data collection, obtaining graph structure data with nodes and edges; Data formatting: converting data into a format suitable for graph convolutional network input, including node feature matrix and adjacency matrix, and standardizing node features to improve training results; Adjacency matrix processing: normalize the adjacency matrix and calculate the inverse square root of the degree matrix of each node's adjacency matrix; Convolution operation, applies convolution operation to each node and its neighborhood to aggregate the information of neighboring nodes; Activation function, use ReLU activation function to increase the nonlinear expression ability of the model; Stacking graph convolutional network layers: stack multiple graph convolutional layers as needed to capture multi-level features of nodes; Output layer: Add an output layer containing a fully connected layer or a pooling layer for the subgraph mining task and train the constructed model; Feature extraction, extracting node and subgraph features from the graph convolutional network output; Subgraph matching, by comparing the extracted features with the target pattern, identifying and extracting subgraphs that match a specific pattern; Deploy the trained graph convolutional network model into the production environment, integrate it into practical applications, use the model for real-time subgraph matching and pattern recognition, and provide data analysis and decision support; S3. Build a relationship model based on the mining results and apply it to specific scenarios for data analysis and insight, where the relationship model includes a graph-based machine learning framework for training and predicting the development and changes of complex relationships between entities; The graph-based machine learning framework in the relational model includes using a graph neural network to model and predict complex relationships between entities, wherein the graph neural network is used to capture dependencies and dynamic changes between entities.

2. The big data relationship mining method based on graph computing according to claim 1 is characterized in that: The data cleaning step includes using data deduplication and noise filtering technology, and the data deduplication uses a local sensitive hashing method, which includes the following steps: Convert each data record into a feature vector; Choose a hash function that suits the characteristics of your data; Using the selected hash function, mapping each of the feature vectors to one or more hash buckets, each of the hash buckets containing a plurality of the feature vectors, and the feature vectors of each hash bucket are relatively close in the hash space; For each new data record, calculate the hash value of the feature vector and map it to the corresponding hash bucket; Performing duplicate detection in the hash bucket, comparing the similarity of all the feature vectors in the same hash bucket, identifying records with a similarity higher than a set threshold as duplicate data, removing the identified duplicate data from the data set, and retaining unique data records; The noise filtering technology includes using a statistical method and a machine learning method, and specifically includes the following steps: The statistical methods include: standard deviation analysis, calculating the standard deviation of each data point and identifying those data points that deviate greatly from the mean; Distribution anomaly detection, using box plots to identify outliers in the data; Apply the Z-Score method, where data points with a Z-Score greater than 3 or less than -3 are considered outliers; The machine learning-based method uses local anomaly factors, specifically: Distance calculation, use Euclidean distance to calculate neighborhood distance, for each data point, calculate its distance to all other data points; Construct the local density, select the k value, and determine the size of the local neighborhood, that is, the k nearest neighbors of each data point; Calculate k distances. For each data point, find the maximum distance among its k nearest neighbors. Calculate the reachable distance. For data points i and j, calculate their reachable distance, which is defined as the larger value between k distances i and j; Local density, calculate the local density of each data point, that is, the density of the data points in its k neighborhoods; Relative density, calculate the relative density of each data point, that is, the ratio of its local density to the local density of the data points in its neighborhood; LOF value, calculate the LOF value of each data point, that is, the average value of its relative density, set the threshold, and set the threshold according to the LOF value. Data points with LOF values ​​exceeding the threshold are considered outliers; Filter outliers to remove or mark these abnormal points from the dataset.

3. The big data relationship mining method based on graph computing according to claim 2 is characterized in that: The entity recognition step includes using a named entity recognition model based on deep learning, wherein the recognition model is a model combining a long short-term memory network and a conditional random field, and the specific method steps are as follows: Obtain an annotated text dataset, which contains annotated entity information; construct a vocabulary and a tag set, wherein the vocabulary maps words in the text to unique index values; The tag set defines entity tag categories; Convert words to vectors using pre-trained word embeddings or custom word embedding layers; Bi-LSTM Network: The input layer uses word embedding as the input of Bi-LSTM. The Bi-LSTM layer uses bidirectional LSTM to capture contextual information. CRF layer: Sequence labeling: Use the CRF layer to label the output of Bi-LSTM and optimize entity boundary recognition; Model training, using the log-likelihood of CRF as the loss function, using the Adam optimization algorithm to train the model, dividing the data into training set and validation set, performing model training, and adjusting hyperparameters to optimize performance.

4. The big data relationship mining method based on graph computing according to claim 3 is characterized in that: The relationship extraction step includes using a deep learning-based relationship extraction model, which uses a pre-trained language model to perform relationship prediction, and includes the following steps: Obtain a large-scale corpus containing entity pairs and relation annotations; Organize the text into a format suitable for model input and divide the dataset into training, validation and test sets; Text tokenization, using BERT’s tokenizer to split the text into vocabulary indices; Generate input and create input samples for each entity pair, including: Token IDs, vocabulary index of text; Attention Masks, indicating which locations should be paid attention to by the model; Token Type IDs, used to distinguish between two different parts of text; Model building, loading pre-trained BERT, and selecting the appropriate BERT variant; BERT layer, extracting contextual information from text; Relation classification head, which adds a fully connected layer on the CLS token output by BERT for relation classification; Define the loss function and use the cross entropy loss function to optimize the classification task; Perform model training and input new text into the trained model to perform relationship prediction and map the relationship categories output by the model back to the actual relationship labels.

5. The big data relationship mining method based on graph computing according to claim 1 is characterized in that: The graph traversal technology includes using a distributed graph computing framework to process large-scale graph data to accelerate the subgraph mining process; Confirm the Spark cluster version and download the GraphX ​​package compatible with the Spark version; Upload the GraphX ​​package to all nodes of the Spark cluster and place it in the specified library directory; Adjust the spark.executor.memory and spark.driver.memory parameters to allocate appropriate memory; Set the spark.cores.max and spark.executor.instances parameters to configure the number of cores and executor instances; Data loading, using Spark's textFile or sequenceFile method to load data from HDFS, S3, or other storage systems; Clean the data, remove invalid or erroneous records, and convert the raw data into vertex and edge formats; Create vertex RDD and edge RDD, ensuring that the data types are consistent with GraphX ​​requirements; Create a vertex RDD and use Spark's parallelize method to create a vertex RDD, assigning a unique ID and corresponding attributes to each vertex; Create edge RDD, use Spark's parallelize method to create edge RDD, define the source vertex, target vertex and attributes of the edge; Build a Graph object, pass the vertex RDD and edge RDD to the Graph constructor, create a Graph object, and set the default vertex attributes for the Graph; Define graph calculation logic: Select appropriate graph algorithms based on business needs and implement algorithm logic; Use GraphX ​​API to call GraphX's pregel and mapVertices APIs for graph calculations; Write user-defined functions to manipulate vertex and edge properties; Distributed graph computing, graph partitioning, and choosing appropriate partitioning strategies; Consider the characteristics of the graph to choose the best partitioning strategy; Parallel computing uses the distributed computing capabilities of GraphX ​​to split graph data into multiple partitions and process them in parallel on the cluster, monitor the progress of tasks and resource usage, and ensure that computing tasks run efficiently.

6. The big data relationship mining method based on graph computing according to claim 1, characterized in that: The data analysis and insight step includes using graph-based visualization technology to present the mined relationships and patterns as interactive graphs to facilitate users to conduct in-depth analysis and understanding.

7. A computer storage medium, characterized in that: The computer storage medium stores computer instructions, which, when called, are used to execute a big data relationship mining method based on graph computing as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Method and system for mining and checking fraud gang relationship in Internet

    CN110413707A

  • Multi-field cross-chapter event causal relationship mining and judging device and method

    CN117453802A