Big data relation mining analysis method based on graph neural network

By constructing data relationship graphs and using graph neural network models and clustering algorithms, the complex relationship mining problem of multi-source heterogeneous data is solved, and efficient and accurate data analysis and business support are achieved.

CN120371890AInactive Publication Date: 2025-07-25CHUZHOU VOCATIONAL & TECHN COLLEGE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510409078.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to effectively reveal complex relationships when processing multi-source heterogeneous data, especially on large-scale dynamic graph datasets, the model convergence speed is slow, overfitting and insufficient generalization capabilities, and the mining results are out of touch with the actual business.

Method used

By constructing a data relationship diagram of multi-source data, using graph embedding technology to map it into low-dimensional vectors, combining graph neural network models for forward propagation and backpropagation training, applying clustering algorithms and association rules mining, combining domain knowledge to build a relational knowledge base and visually display it.

Benefits of technology

It improves the efficiency and accuracy of data analysis, enhances the prediction accuracy and robustness of the model, can identify complex patterns and high-order associations, directly serves specific business needs, and improves the interpretability and operability of data analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371890A_ABST
    Figure CN120371890A_ABST
Patent Text Reader

Abstract

The invention discloses a big data relation mining analysis method based on a graph neural network, which comprises the following steps: S1, acquiring multi-source data, extracting data features, constructing a data relation graph, and mapping the data relation graph into a low-dimensional vector as a basic data structure; s2, inputting a low-dimensional vector through the graph neural network model, outputting a prediction result, calculating an error by using a loss function based on a real relation label, adjusting parameters of the graph neural network model, and obtaining a target data model; s3, obtaining a data relationship prediction result through the target data model, mining a data potential relationship, and mining a data relationship result through the data potential relationship; and S4, combining a data relationship result with domain knowledge and business rules, constructing a relationship knowledge base, and displaying the data relationship through a visualization technology. According to the method, efficient and accurate big data relationship mining is realized, and the accuracy and practicability of potential relationship discovery in a complex data environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network data relationship processing, and particularly relates to a big data relationship mining and analysis method based on a graph neural network. Background Art

[0002] With the rapid development of information technology and the advent of the big data era, the amount of data has shown an explosive growth. It has become increasingly difficult to mine valuable information from numerous data, especially when dealing with multi-source heterogeneous data. Traditional analysis methods are difficult to effectively reveal the complex relationships between data. Therefore, there is an urgent need for a new technical solution to solve these problems in order to achieve efficient mining and analysis of potential relationships in big data. The big data relationship mining and analysis method based on a graph neural network is precisely proposed to address this challenge, aiming to provide a more effective data analysis means by combining the advantages of graph theory and deep learning.

[0003] Traditionally, the relationship mining of structured data mainly relies on statistical methods and machine learning algorithms. Although these methods can identify some simple linear or non-linear relationships, their mining effect for complex and non-explicit relationships is limited. In addition, they can often only process specific types of input data and have insufficient support for semi-structured and unstructured data. At the same time, traditional methods usually lack an understanding of the topological structure between data, are less efficient when dealing with large-scale data sets, and are difficult to adapt to the constantly changing data environment. Nevertheless, traditional methods have certain advantages in terms of computational resource consumption and model interpretability.

[0004] Existing technologies have tried to apply graph theory to relationship mining, such as in the fields of social network analysis and recommendation systems, etc. However, most existing technologies focus on simple feature extraction on static graphs and fail to fully utilize the powerful expressive ability of graph neural networks. At the same time, existing technologies do not deeply mine node attribute information and the deep topological structure of the graph, and there are problems of slow convergence speed and overfitting when dealing with large-scale dynamic graphs, resulting in low model generalization ability and prediction accuracy. In addition, existing technologies rarely combine mining results with domain knowledge and cannot directly provide support for specific business decisions.

[0005] The big data relationship mining and analysis method based on a graph neural network proposed by the present invention not only solves the problem of insufficient multi-source data processing ability in existing technologies, but also improves the mining accuracy of complex data relationships. Summary of the Invention

[0006] A big data relationship mining and analysis method based on a graph neural network according to the present invention includes:

[0007] S1. Obtain multi-source data, extract data features, construct a data relationship graph using distance metrics and information theory, and map the data relationship graph into a low-dimensional vector through graph embedding technology as the basic data structure;

[0008] S2. Input the low-dimensional vector through a graph neural network model, perform forward propagation in the graph neural network model, output the prediction result, calculate the error using a loss function based on the true relationship label, backpropagate to adjust the parameters of the graph neural network model, and use regularization technology to prevent overfitting. Train until the graph neural network model converges to obtain the target data model;

[0009] S3. Obtain the data relationship prediction result through the target data model, process the prediction result using clustering algorithms and association rule mining algorithms, mine the potential data relationships, and mine the data relationship results through the potential data relationships;

[0010] S4. Combine the data relationship results with domain knowledge and business rules to construct a relationship knowledge base, and display the data relationships through visualization technology.

[0011] Preferably, the construction of the S1 data relationship graph includes:

[0012] S1.1. Obtain the structured data, semi-structured, and unstructured data of the multi-source data, extract the multi-source data features, and obtain the numerical data features and discrete data features of the multi-source data;

[0013] S1.2. Calculate the similarity between numerical data features through distance metrics, and calculate the dependence degree between discrete feature data through information theory;

[0014] S1.3. Obtain the measurement results of numerical data features and discrete data features, and construct an initial data relationship graph.

[0015] Preferably, in S1.2, the similarity between numerical data features is calculated by the Euclidean distance of distance metrics, and the dependence degree between discrete feature data is calculated by the mutual information of information theory, specifically including:

[0016] The numerical data features include X = (x1, x2,..., x n ) and Y = (y1, y2,..., y n ), and the measurement formula for the similarity S n is: where n is the number of dimensions, x i and y i are the i-th numerical data features respectively; the similarity S n is the similarity of numerical data features, and the closer the similarity S n is to 1, the more similar the numerical data features of X and Y are; the similarity Sn The closer it is to 0, the lower the similarity of the numerical data features of X and Y;

[0017] Discrete feature data includes A and B, where A has m category values and B has k category values; the variables A and B are used to construct a joint probability distribution matrix P(A, B), and the mutual trust M between A and B is calculated. I , the formula is: Where P(a i ,b j ) is A with a value of a i and B takes the value b j The probability value of the discrete feature data A and B is I The larger the value is, the higher the degree of dependence between the discrete feature data A and B is; conversely, the lower the degree of dependence between the discrete feature data A and B is.

[0018] Preferably, the construction of the initial data relationship graph in S1.3 is carried out by obtaining the measurement results of the numerical data features and the measurement results of the discrete data features, and according to a pre-set similarity threshold, the numerical data points with higher similarity under the distance measurement and the discrete features with strong dependency relationships are connected and incorporated into the graph structure to form a data relationship graph.

[0019] Preferably, the graph embedding technology is used in S1 to map the data relationship graph into a low-dimensional vector as the basic data structure, specifically including:

[0020] The node-level attribute information of the data relationship graph is extracted, including but not limited to the number of node connections, the connection direction and the relative position of the node, and the node attribute information is converted into a high-dimensional feature vector; the adjacent node matrix is constructed through the topological structure using the weights of the edges between the nodes of the data relationship graph, the types of the edges and the path lengths between the nodes; the high-dimensional feature vector and the adjacent node matrix are used as input through the graph embedding algorithm, and multiple random walk sampling or matrix decomposition is performed on the nodes of the data relationship graph to mine the potential semantic information and structure of the nodes in the data relationship graph to obtain low-dimensional vectors of the nodes.

[0021] Preferably, the updating of the graph neural network model parameters in S2 specifically includes:

[0022] The prediction results are output through the graph neural network model The error L is calculated by the mean square error loss function s , the formula is: where y iis the i-th true relationship label, and N is the number of samples; adjusting the parameters of the graph neural network model through backpropagation based on the error value; the backpropagation calculates the gradients of each layer of parameters in turn using the chain rule according to the partial derivatives of the output layer parameters with respect to the loss function, starting from the output layer, and gradually calculates the contribution degree of each parameter to the loss function in the direction of the input layer to obtain the gradient values of each layer of parameters, and updates the parameters of the graph neural network model through the gradient descent algorithm.

[0023] Preferably, the target data model is trained to convergence of the graph neural network model through the gradient descent algorithm and regularization technology, specifically including:

[0024] Updating the parameters of the graph neural network model through the gradient descent algorithm, the formula is: where θ n is the updated parameter of the graph neural network model, θ o is the original parameter of the graph neural network model, α is the learning rate, is the gradient change rate of the parameters of the graph neural network model; iteratively updating the parameters through backpropagation, and using L1 or L2 regularization technology to constrain the parameters until the graph neural network model converges to obtain the target data model.

[0025] Preferably, in S3, the prediction results are processed using the clustering algorithm and the association rule mining algorithm to mine the potential relationships of the data, specifically including:

[0026] After obtaining the data relationship prediction results output by the target data model, the data in the prediction results are divided into different clusters according to the similarity between the data through the clustering algorithm, and the data with close distances and similar features are grouped into the same cluster, presenting the distribution pattern and potential grouping of the data; through the association rule mining algorithm, search for each cluster and the data within the cluster, and identify the data combinations that frequently co-occur between different clusters or within the cluster, and identify the associated data with high support and confidence as the potential relationship clusters.

[0027] Preferably, in S3, the data relationship results are mined using the potential relationships of the data, specifically including:

[0028] Determine the tightness and interaction mode between data elements according to the potential relationship clusters, construct a local relationship network framework, determine the key data elements with bridging effects by comparing the connection points and interaction frequencies between different potential relationship clusters as the key nodes of the global relationship network, and connect each local network in series to form a complete data relationship map; through the data relationship map, use the path analysis algorithm to trace the propagation path and dependence path between data elements, quantify the weights and credibility on different paths, and mine the data relationship results.

[0029] Preferably, in S4, the data relationship result is combined with domain knowledge and business rules to construct a relationship knowledge base, specifically including:

[0030] Obtain the data relationship result, extract concepts, principles, and empirical rules closely related to the domain knowledge from the domain knowledge, integrate them into the data relationship, and clarify the meaning and role of data nodes in the domain logic; according to the domain business rules, screen the data relationship, ensure the data relationship, and organize the integrated data relationship in a structured manner to construct a relationship knowledge base.

[0031] Compared with the prior art, the technical solution of the present application has the following technical effects:

[0032] The present invention constructs a data relationship graph and uses graph embedding technology to transform complex data structures into low-dimensional vectors, solving the problem that traditional methods are difficult to effectively capture deep-level associations between data when dealing with multi-source heterogeneous data. This solution can not only handle various types of data features such as numerical and discrete types, but also accurately quantify the similarity and dependence degree between different data points through distance metrics and information theory. Thus, we obtain a more comprehensive and accurate data representation form, enabling subsequent analysis and mining work to be carried out on a more optimized basic data structure, thereby improving the efficiency and accuracy of data analysis.

[0033] The present invention introduces a graph neural network model to perform forward propagation on the mapped low-dimensional vectors, and combines the backpropagation algorithm and regularization technology to train the model parameters until convergence, solving the problems of overfitting and insufficient generalization ability of the model in the prior art. Utilizing the powerful expression ability and adaptive adjustment mechanism of deep learning, it ensures that the model can converge quickly and stably on large-scale dynamic graph datasets, obtaining a target data model with high prediction accuracy and strong robustness, and providing more reliable prediction results.

[0034] The present invention mines the potential relationships of data by applying clustering algorithms and association rule mining algorithms to the prediction results, solving the problem of difficult recognition of complex patterns and high-order associations in the prior art. Classify them into different clusters according to the similarity between data, and further search for frequently co-occurring data combinations, revealing potential relationships with high support and confidence hidden in the data, and extracting valuable patterns and rules from a large amount of data, providing a scientific basis for decision-making, enhancing the understanding of the internal structure of data, and promoting the development of data-driven businesses.

[0035] By combining the mined data relationships with domain knowledge and business rules, the present invention constructs a relationship knowledge base and uses visualization technology to display the data relationships, solves the problem in the prior art that the mining results are disconnected from the actual application scenarios, concretizes the abstract data analysis results, enables them to directly serve the business needs of specific fields, improves the interpretability and operability of the data analysis results, and can also help users more intuitively understand the meaning behind the data, promotes the sharing and transmission of knowledge, and realizes the effective transformation from data to wisdom.

[0036] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, so as to be implemented in accordance with the content of the specification, and in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following takes the preferred embodiments of the present application and combines the accompanying drawings to describe in detail as follows.

[0037] According to the following detailed description of the specific embodiments of the present application in conjunction with the accompanying drawings, those skilled in the art will be more clear about the above and other purposes, advantages and features of the present application. Description of the Drawings

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to actual scale.

[0039] Figure 1 It is a flowchart of a big data relationship mining and analysis method based on graph neural network of the present invention;

[0040] Figure 2 It is a flowchart of constructing a data relationship graph of a big data relationship mining and analysis method based on graph neural network of the present invention;

[0041] Figure 3 It is a flowchart of the clustering result of a big data relationship mining and analysis method based on graph neural network of the present invention.

[0042] Figure 4 It is a flowchart of the potential relationship cluster of a big data relationship mining and analysis method based on graph neural network of the present invention; Detailed Embodiments

[0043] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some, but not all, of the embodiments of this application. In the following description, specific details such as specific configurations and components are provided only to assist in a comprehensive understanding of the embodiments of this application. Therefore, those skilled in the art should clearly understand that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Additionally, descriptions of known functions and configurations are omitted for clarity and conciseness in the embodiments.

[0044] It should be understood that the term "one embodiment" or "this embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, the appearances of the term "one embodiment" or "this embodiment" throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner.

[0045] In addition, this application may repeat reference numerals and / or letters in different instances. This repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or arrangements discussed.

[0046] The term "and / or" in this document is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, B exists alone, and both A and B exist simultaneously. The term " / and" in this document describes another association relationship of associated objects, indicating that two relationships can exist. For example, A / and B can represent: A exists alone, and both A and B exist. Additionally, the character " / " in this document generally indicates that the associated objects before and after are in an "or" relationship.

[0047] The term "at least one" in this document is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, at least one of A and B can represent: A exists alone, both A and B exist simultaneously, and B exists alone.

[0048] It should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise," "include," or any other variant thereof are intended to cover non-exclusive inclusion.

[0049] Embodiment 1

[0050] This embodiment mainly describes a big data relationship mining and analysis method based on graph neural network, as Figure 1 shown, including the following steps:

[0051] S1. Obtain multi-source data, extract data features, construct a data relationship graph using distance metrics and information theory, and map the data relationship graph to a low-dimensional vector through graph embedding technology as the basic data structure;

[0052] S2. Input the low-dimensional vector through the graph neural network model, perform forward propagation in the graph neural network model, output the prediction result, calculate the error using the loss function based on the true relationship label, backpropagate to adjust the parameters of the graph neural network model, and use regularization technology to prevent overfitting, and train until the graph neural network model converges to obtain the target data model;

[0053] S3. Obtain the data relationship prediction result through the target data model, process the prediction result using clustering algorithms and association rule mining algorithms, mine the potential data relationships, and mine the data relationship result through the potential data relationships;

[0054] S4. Combine the data relationship result with domain knowledge and business rules to construct a relationship knowledge base, and display the data relationship through visualization technology.

[0055] Furthermore, as Figure 2 shown, the construction of the S1 data relationship graph includes the following steps:

[0056] S1.1. Obtain the structured data, semi-structured and unstructured data of multi-source data, extract the multi-source data features, and obtain the numerical data features and discrete data features of multi-source data;

[0057] S1.2. Calculate the similarity between numerical data features through distance metrics, and calculate the dependence degree between discrete feature data through information theory;

[0058] S1.3. Obtain the measurement results of numerical data features and discrete data features, and construct an initial data relationship graph.

[0059] Furthermore, in S1.2, the similarity between numerical data features is calculated by the Euclidean distance of distance metrics, and the dependence degree between discrete feature data is calculated by the mutual information of information theory, specifically including:

[0060] The numerical data features include X = (x1, x2,..., x n ) and Y = (y1, y2,..., y n ), and the measurement formula of similarity S n is: where n is the number of dimensions, x i and yi They are the i-th numerical data features respectively; similarity S n is the similarity of numerical data features, similarity S n The closer it is to 1, the more similar the numerical data features of X and Y are; similarity S n The closer it is to 0, the lower the similarity of the numerical data features of X and Y;

[0061] The discrete feature data includes A and B. A has m category values and B has k category values; a joint probability distribution matrix P(A,B) is constructed using variables A and B, and the mutual trust degree M between A and B is calculated I , and the formula is: where P(a i ,b j ) is the probability value when A takes the value of a i and B takes the value of b j ; the mutual trust degree M of the discrete feature data A and B I The larger it is, the higher the dependence degree of the discrete feature data A and B. On the contrary, the lower the dependence degree of the discrete feature data A and B.

[0062] Furthermore, the construction of the initial data relationship graph in S1.3 is based on the measurement results of the obtained numerical data features and the measurement results of the discrete data features. According to the preset similarity threshold, the numerical data points with higher similarity under distance measurement and the discrete features with strong dependence relationships are connected and incorporated into the graph structure to form a data relationship graph;

[0063] The preset similarity threshold is obtained by using the sequential forward selection method or the backward selection method, and the classifier is used to select the optimal feature subset for data relationship mining and analysis;

[0064] The sequential forward selection method starts from an empty feature subset and each time selects a feature that can most significantly improve the performance of the classifier and adds it to the subset. By continuously iterating this process, a subset containing the optimal feature combination is obtained, which can dynamically select features according to the characteristics of the data itself and the feedback of the classifier;

[0065] The backward selection method starts from the full set containing all features and each time removes a feature that has the least impact on the performance of the classifier, continuously deletes unimportant features, and obtains an optimal feature subset;

[0066] After selecting the optimal feature subset using a classifier, the threshold is determined using cross-validation, the dataset is divided into a training set and a validation set, the classifier is used to train on the training set, and the validation set is predicted using different similarity thresholds. Observe the classification results on the validation set, plot the precision-recall curve (PR curve) at different thresholds, and select the threshold that maximizes the F1 value on the curve as the similarity threshold.

[0067] Further, in S1, the graph embedding technology is used to map the data relationship graph into a low-dimensional vector as the basic data structure, specifically including:

[0068] Extract the node-level attribute information of the data relationship graph, including but not limited to the number of node connections, connection directions, and relative node positions, and convert the node attribute information into a high-dimensional feature vector; through the topological structure, construct an adjacency node matrix using the weights of the edges between the nodes, the types of edges, and the path lengths between the nodes in the data relationship graph; use the high-dimensional feature vector and the adjacency node matrix as inputs through the graph embedding algorithm, perform multiple random walk samplings or matrix decompositions on the nodes of the data relationship graph, mine the latent semantic information and structure of the nodes in the data relationship graph, and obtain the low-dimensional vectors of the nodes.

[0069] Further, the update of the parameters of the graph neural network model in S2 specifically includes:

[0070] Output the prediction result through the graph neural network model Calculate the error L through the mean squared error loss function s , the formula is: where y i is the i-th true relationship label, and N is the number of samples; through the error value, backpropagation is performed to adjust the parameters of the graph neural network model; backpropagation calculates the gradient of each layer of parameters in turn according to the partial derivative of the loss function with respect to the output layer parameters using the chain rule, starting from the output layer, gradually calculating the contribution degree of each parameter to the loss function in the direction of the input layer, obtaining the gradient values of each layer of parameters, and updating the parameters of the graph neural network model through the gradient descent algorithm.

[0071] Further, the target data model is trained to convergence of the graph neural network model through the gradient descent algorithm and regularization technology, specifically including:

[0072] Update the parameters of the graph neural network model through the gradient descent algorithm, the formula is: where, θ n is the updated parameter of the graph neural network model, θ o is the original parameter of the graph neural network model, and α is the learning rate. is the gradient change rate of the parameters of the graph neural network model; the parameters are iteratively updated through backpropagation, and the L1 or L2 regularization technique is used to constrain the parameters until the graph neural network model converges to obtain the target data model.

[0073] Further, in S3, the prediction results are processed using clustering algorithms and association rule mining algorithms to mine the potential relationships in the data, specifically including:

[0074] After obtaining the data relationship prediction results output by the target data model, the data in the prediction results are divided into different clusters according to the similarity between the data through the clustering algorithm, and the data with close distances and similar features are grouped into the same cluster, presenting the distribution pattern and potential grouping of the data; through the association rule mining algorithm, each cluster and the data within the cluster are searched, and the data combinations that frequently co-occur between different clusters or within the cluster are identified, and the associated data with high support and confidence are identified as potential relationship clusters.

[0075] Further, in S3, the potential relationships in the data are used to mine the data relationship results, specifically including:

[0076] According to the potential relationship clusters, the tightness and interaction mode between data elements are determined, a local relationship network framework is constructed, and by comparing the connection points and interaction frequencies between different potential relationship clusters, the key data elements with bridging effects are determined as the key nodes of the global relationship network, and each local network is connected in series to form a complete data relationship map; through the data relationship map, the path analysis algorithm is used to trace the propagation path and dependence path between data elements, quantify the weights and credibility on different paths, and mine the data relationship results.

[0077] Further, in S4, the data relationship results are combined with domain knowledge and business rules to construct a relationship knowledge base, specifically including:

[0078] Obtain the data relationship results, extract the concepts, principles and empirical rules closely related to the data relationship from the domain knowledge, integrate them into the data relationship, and clarify the meaning and role of the data nodes in the domain logic; according to the domain business rules, screen the data relationship to ensure the data relationship, and organize the integrated data relationship in a structured manner to construct a relationship knowledge base.

[0079] In this embodiment, different measurement methods are respectively used for numerical and discrete data features to construct an initial data relationship graph, and advanced graph embedding technology and regularization strategies are used to train the graph neural network model, so that the potential connections between data can be captured more accurately. By also integrating the mined data relationships with domain knowledge, a relationship knowledge base is constructed, thus providing high-value decision support for actual business.

[0080] Embodiment 2

[0081] In this embodiment, the data relationship prediction result is obtained through the target data model, and the prediction result is processed by using the clustering algorithm and the association rule mining algorithm to mine the potential data relationship, and the data relationship result is mined through the potential data relationship, specifically including;

[0082] As Figure 3 shown, in the big data relationship mining and analysis process based on the graph neural network, after successfully obtaining the data relationship prediction result output by the target data model, the potential data relationship is mined;

[0083] Use the clustering algorithm, such as the K-Means clustering algorithm based on Euclidean distance, to finely divide the data in the prediction result. Assume that the data point set is L = (l1, l2,..., l n ), where each data point l i has an r-dimensional feature vector, that is, l i = (l i1 , l i2 ,..., l im ). For the given number of clusters q, randomly initialize q cluster centers μ1, μ2,..., μ q , where each cluster center μ j = (μ j1 , μ j2 ,..., μ jm ). Calculate the Euclidean distance from each data point l i to each cluster center μ j Assign the data point to the cluster represented by the nearest cluster center, that is, the data point l i is assigned to the cluster C j if and only if Recalculate the new cluster center of each cluster, and the formula is: where |C j | is the number of data points in the cluster C j ; continuously repeat the above steps of assigning and updating the cluster center until the cluster center no longer changes significantly or reaches the preset maximum number of iterations T. According to the similarity between data, group the data with close distances and similar features into the same cluster, clearly presenting the distribution pattern and potential grouping of the data.

[0084] As Figure 4 shown, after completing the clustering, use the Apriori algorithm to deeply search each cluster and the data within the cluster; by scanning the clustered data set, count the occurrence frequency of each item set (the item set is the clustered cluster or the feature in the cluster), that is, the support count. For an item set I, its support where c(I) is the number of times the item set I appears in the data set, and N J is the total number of records in the data set; if the support of an item set is greater than a pre-set minimum support threshold, then the item set is considered a frequent item set, and then association rules are generated based on the frequent item set. For an association rule a → b, its confidence The strength of the association rule is measured by calculating the confidence, and the association rules with a confidence greater than the minimum confidence threshold are retained. The data combinations involved in the association rules are the clusters with potential relationships, and the identified association data with high support and confidence form potential relationship clusters.

[0085] This embodiment details the mining of data relationship results through data potential relationships, providing a key foundation for subsequent further analysis and utilization of complex relationships in big data, strongly supporting the mining and analysis of big data relationships based on graph neural networks, so as to more accurately reveal the value information and laws hidden behind the data, and providing strong data-driven support for decision-making and business optimization in various fields.

[0086] The above are only the preferred embodiments of the present invention, and it does not limit the protection scope of the present invention. For those skilled in the art, the present invention can have various changes and modifications; within the spirit and principle of the present invention, through conventional substitutions or the ability to achieve the same function, without departing from the principle and spirit of the present invention, changes, modifications, substitutions, integrations, and parameter changes to these embodiments all fall within the protection scope of the present invention.

Claims

1. A method for mining and analyzing big data relationships based on graph neural networks, characterized in that Including: S1. Obtain multi-source data, extract data features, construct a data relationship graph using distance metrics and information theory, and map the data relationship graph into a low-dimensional vector through graph embedding technology as the basic data structure; S2. Input the low-dimensional vector through a graph neural network model, perform forward propagation in the graph neural network model, output the prediction result, calculate the error using a loss function based on the true relationship label, backpropagate to adjust the parameters of the graph neural network model, and use regularization technology to prevent overfitting, and train until the graph neural network model converges to obtain the target data model; S3. Obtain the data relationship prediction result through the target data model, process the prediction result using clustering algorithms and association rule mining algorithms, mine the potential data relationships, and mine the data relationship result through the potential data relationships; S4. Combine the data relationship result with domain knowledge and business rules to construct a relationship knowledge base, and display the data relationship through visualization technology.

2. The method for mining and analyzing big data relationships based on a graph neural network according to claim 1, wherein The construction of the data relationship graph in S1 includes: S1.

1. Obtain the structured data, semi-structured and unstructured data of multi-source data, convert the semi-structured and unstructured data into a structured form, extract the data features of the structured form, and obtain numerical and discrete data features; S1.

2. Calculate the similarity between numerical data features through distance metrics, and calculate the dependence degree between discrete feature data through mutual information of information theory; S1.

3. Obtain the measurement results of numerical data features and discrete data features, and construct an initial data relationship graph.

3. A method for mining and analyzing big data relationships based on a graph neural network according to claim 1 or 2, characterized in that, In S1.2, the similarity between numerical data features is calculated through the Euclidean distance of distance metrics, and the dependence degree between discrete feature data is calculated through mutual information of information theory, specifically including: The numerical data features include X = (x1, x2, …, x n ) and Y = (y1, y2, …, yn), and the similarity S n is measured by the formula: where n is the number of dimensions, x i and y i are the i-th numerical data features respectively; the similarity S n is the similarity of the numerical data features. The closer the similarity S n is to 1, the more similar the numerical data features of X and Y are; the closer the similarity S n is to 0, the lower the similarity of the numerical data features of X and Y is; The discrete feature data includes A and B. A has m categories and B has k categories. The joint probability distribution matrix P(A, B) is constructed using variables A and B, and the mutual trust degree M between A and B is calculated. I , and the formula is: where P(a i , b j ) is the probability value when A takes the value of a i and B takes the value of b j . The larger the mutual trust degree M I of the discrete feature data A and B, the higher the degree of dependence between the discrete feature data A and B. Conversely, the lower the degree of dependence between the discrete feature data A and B.

4. A method for mining and analyzing big data relationships based on a graph neural network according to claim 1 or 2, characterized in that In S1.3, the construction of the initial data relationship graph is based on the measurement results of the obtained numerical data features and discrete data features. According to a preset similarity threshold, connect the numerical data points with higher similarity under distance metrics and the discrete features with strong dependence relationships, and incorporate them into the graph structure to form a data relationship graph.

5. A big data relationship mining and analysis method based on a graph neural network according to claim 1, characterized in that, In S1, the data relationship graph is mapped into a low-dimensional vector as the basic data structure using graph embedding technology, specifically including: Extract the node-level attribute information of the data relationship graph, including but not limited to the number of node connections, connection directions, and relative node positions, and convert the node attribute information into a high-dimensional feature vector; construct an adjacency node matrix through the topological structure using the weights of the edges between the nodes of the data relationship graph, the types of edges, and the path lengths between the nodes; use the graph embedding algorithm to take the high-dimensional feature vector and the adjacency node matrix as inputs, perform multiple random walk samplings or matrix decompositions on the nodes of the data relationship graph, mine the potential semantic information and structure of the nodes in the data relationship graph, and obtain the low-dimensional vector of the nodes.

6. A method for mining and analyzing big data relationships based on a graph neural network according to claim 1, characterized in that, The update of the parameters of the graph neural network model in S2 specifically includes: The prediction result is output through the graph neural network model The error L is calculated through the mean square error loss function s , and the formula is: where y i is the i-th true relationship label, and N is the number of samples; the parameters of the graph neural network model are adjusted through backpropagation using the error value; in the backpropagation, according to the partial derivative of the loss function with respect to the parameters of the output layer, the gradients of each layer of parameters are calculated in turn using the chain rule. Starting from the output layer, the contribution degree of each parameter to the loss function is calculated step by step in the direction of the input layer to obtain the gradient values of each layer of parameters, and the parameters of the graph neural network model are updated through the gradient descent algorithm.

7. A method for mining and analyzing big data relationships based on a graph neural network according to claim 6, characterized in that, The target data model is trained until the graph neural network model converges through the gradient descent algorithm and regularization technology, specifically including: The parameters of the graph neural network model are updated by the gradient descent algorithm, and the formula is as follows: where θ n is the parameter of the updated graph neural network model, θ o is the parameter of the original graph neural network model, α is the learning rate, is the gradient change rate of the parameters of the graph neural network model; the parameters are updated iteratively by backpropagation, and the L1 or L2 regularization technique is used to constrain the parameters until the graph neural network model converges to obtain the target data model.

8. A method for mining and analyzing big data relationships based on a graph neural network according to claim 1, characterized in that In S3, the prediction results are processed using clustering algorithms and association rule mining algorithms to mine the potential relationships in the data, specifically including: After obtaining the data relationship prediction results output by the target data model, the data in the prediction results are divided into different clusters according to the similarity between the data through a clustering algorithm. The data with close distances and similar features are grouped into the same cluster, presenting the distribution pattern and potential grouping of the data. Through the association rule mining algorithm, each cluster and the data within the cluster are searched for data combinations that frequently co-occur between different clusters or within the cluster, and the associated data with high support and confidence are identified as potential relationship clusters.

9. A method for mining and analyzing big data relationships based on a graph neural network according to claim 1 or 8, characterized in that, In S3, the data relationship results are mined using the potential relationships in the data, specifically including: Based on the potential relationship clusters, the degree of closeness and the interaction mode between data elements are determined, and a local relationship network framework is constructed. By comparing the connection points and interaction frequencies between different potential relationship clusters, the key data elements with bridging effects are determined as the key nodes of the global relationship network, and each local network is connected in series to form a complete data relationship map. Through the data relationship map, the path analysis algorithm is used to trace the propagation paths and dependence paths between data elements, quantify the weights and credibility on different paths, and mine the data relationship results.

10. A method for mining and analyzing big data relationships based on a graph neural network according to claim 1, characterized in that, In S4, the data relationship results are combined with domain knowledge and business rules to construct a relationship knowledge base, specifically including: The data relationship results are obtained, and the concepts, principles, and empirical rules closely related to the data relationship are extracted from the domain knowledge and incorporated into the data relationship to clarify the meaning and role of the data nodes in the domain logic. According to the domain business rules, the data relationships are screened to ensure the data relationships, and the integrated data relationships are organized in a structured manner to construct a relationship knowledge base.

Citation Information

Cited By

  • Business data association analysis method and system based on graph neural network

    CN121352156A