Depth attribute graph clustering method based on multi-information comparison
By fusing node attributes and graph structure information, and adopting graph neural networks and multi-information comparative learning strategies, the problems of insufficient information utilization and insufficient feature representation in existing attribute graph clustering methods are solved, achieving more efficient and accurate node clustering.
Patent Information
- Application Number
- CN202510821366.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-16
AI Technical Summary
Existing attribute graph clustering methods have problems such as insufficient information utilization, limited feature representation capabilities, lack of effective comparative learning mechanisms and low computational efficiency, which lead to one-sided clustering results, unclear boundaries and difficulty in meeting real-time requirements.
By fusing node attribute information and graph structure information, a graph neural network is used to learn the high-order feature representation of nodes, and a multi-information comparative learning strategy is designed to enhance the model's ability to identify node differences. K-Means or DBSCAN clustering algorithms are used for node division.
It improves the accuracy and efficiency of attribute graph clustering, can more comprehensively characterize node characteristics, enhance the ability to identify node differences, improve the clarity of cluster boundaries, and meet real-time requirements.
Smart Images

Figure CN120655952A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data mining and machine learning, and in particular to a deep attribute graph clustering method based on multi-information comparison. Background Art
[0002] In the era of big data, attribute graphs, as a powerful data structure that can simultaneously represent the attribute characteristics of entities (nodes) and the relationships between entities (edges), are widely used in various practical scenarios. For example, in social networks, users can be represented as nodes, with user attributes such as age and interests constituting node attributes, and friendships between users constituting edges. In biological networks, proteins can be represented as nodes, with their functional characteristics constituting node attributes, and interactions between proteins constituting edges.
[0003] The main problems of existing attribute graph clustering methods are as follows: Insufficient information utilization: Most methods only consider a single dimension of node attribute information or graph structure information, and fail to effectively integrate the two types of information, resulting in one-sided clustering results. For example, clustering based only on node attributes will ignore the structural relationship between nodes, while clustering based only on graph structure will lose the attribute characteristics of nodes. Limited feature representation capabilities: Traditional methods find it difficult to learn the high-order semantic features and complex nonlinear relationships of nodes, resulting in insufficient ability to identify node differences, especially when processing large-scale, high-dimensional attribute graph data. Lack of effective comparative learning mechanism: Existing methods cannot enhance the model's sensitivity to node differences by explicitly comparing the feature representations of different nodes, resulting in unclear clustering boundaries and prone to erroneous clustering results. Low computational efficiency: Some methods have high computational complexity when processing large-scale attribute graphs, making it difficult to meet the real-time requirements of practical applications.
[0004] Therefore, there is an urgent need for a deep attribute graph clustering method that can fully integrate multi-source information, enhance the ability to identify node differences and improve clustering accuracy. Summary of the Invention
[0005] The purpose of the present invention is to provide a deep attribute graph clustering method based on multi-information comparison. By fusing node attribute information and graph structure information and designing a multi-information comparison learning strategy, the model's ability to recognize node differences is enhanced, thereby improving the accuracy and efficiency of attribute graph clustering.
[0006] The technical solutions of the present invention are as follows:
[0007] A deep attribute graph clustering method based on multi-information comparison includes the following steps:
[0008] Acquire social attribute graph data, and clean and standardize node attribute information and edge relationship information in the social attribute graph data;
[0009] Fuse the node attribute information and the graph structure information to obtain the fused node feature vector;
[0010] Using the fused node feature vectors to construct a deep attribute graph model, and using a graph neural network to learn high-order feature representations of nodes;
[0011] Design a multi-information comparative learning strategy to enhance the deep attribute graph model's ability to identify node differences by comparing the feature representations of different nodes;
[0012] Cluster analysis is performed using the learned node feature representation to divide the nodes into different clusters to obtain clustering results; the clustering results include multiple set interest groups.
[0013] Furthermore, the node attribute information and edge relationship information are cleaned and standardized, specifically:
[0014] For node attribute information, if there are missing values, use mean filling, median filling, or missing value prediction methods based on machine learning to fill them; for edge relationship information, check whether there are abnormal connections. If there are abnormal connections, delete duplicate connections and self-connections, and correct the directional relationships of incorrect connections.
[0015] Furthermore, the node attribute information and the graph structure information are integrated as follows:
[0016] The node attribute information is represented as a vector form, and the graph embedding technology is used to map the nodes in the graph to a low-dimensional vector space to obtain the node's structural embedding vector, and then the node's attribute vector and structural embedding vector are spliced.
[0017] Furthermore, the graph embedding technology is DeepWalk or Node2Vec method.
[0018] Furthermore, the fused node feature vectors are used to construct a deep attribute graph model, and the high-order feature representation of the nodes is learned using a graph neural network. Specifically:
[0019] The fused node feature vector and the adjacency matrix of the attribute graph are used as the input of the graph neural network. The graph neural network contains two network layers. During the calculation process of each network layer, the features of adjacent nodes are weighted aggregated and nonlinearly transformed according to the connection relationship between the nodes, and the feature representation of the nodes is updated to learn the high-order feature representation of the nodes; wherein, the graph neural network is a graph convolutional network or a graph attention network.
[0020] Furthermore, the design of the multi-information comparative learning strategy specifically includes:
[0021] Construct positive sample pairs and negative sample pairs, where the positive sample pairs are node pairs with dense connections and similar attributes in the graph structure, and the negative sample pairs are node pairs with sparse connections and large attribute differences in the graph structure; define a contrastive learning loss function, and by minimizing the contrastive learning loss function, make the feature representations of the positive sample pairs closer in the feature space, and the feature representations of the negative sample pairs farther apart in the feature space.
[0022] Furthermore, cluster analysis is performed using the learned node feature representation to divide the nodes into different clusters. The clustering results are as follows:
[0023] Calculate the distance or similarity between the learned node feature vectors, and divide the nodes into different clusters according to the clustering rules of the set clustering algorithm.
[0024] Furthermore, the clustering algorithm used in the cluster analysis is a K-Means clustering algorithm or a DBSCAN clustering algorithm.
[0025] Compared with the prior art, the present invention has the following advantages:
[0026] The present invention fully utilizes the multi-dimensional characteristics of data by fusing node attribute information and graph structure information, and can more comprehensively characterize the characteristics of nodes, thereby improving the accuracy of clustering.
[0027] Using graph neural networks to learn high-order feature representations of nodes can capture the complex relationships and nonlinear features between nodes and enhance the model's ability to identify node differences.
[0028] A multi-information comparative learning strategy is designed to explicitly compare the feature representations of different nodes, enabling the model to better distinguish similar nodes from dissimilar nodes and improve the clarity of cluster boundaries.
[0029] The present invention is applicable to attribute graph data in the social field. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings illustrate various embodiments generally by way of example and not limitation, and together with the description and claims, serve to explain embodiments of the invention. Where appropriate, the same reference numerals are used throughout the drawings to refer to the same or similar parts. Such embodiments are illustrative and are not intended to be exhaustive or exclusive of the embodiments of the present apparatus or method.
[0031] Figure 1 A schematic flow chart of the method of the present invention is shown. DETAILED DESCRIPTION
[0032] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0033] like Figure 1 As shown, an embodiment of the present invention provides a deep attribute graph clustering method based on multi-information comparison, including:
[0034] Acquiring and preprocessing data
[0035] In the implementation of this invention, we first need to obtain attribute graph data from social networks. This data comes from various fields. After obtaining the attribute graph data, we need to clean and standardize the node attribute information and edge relationship information in the attribute graph.
[0036] Node attribute information processing: In real-world data, node attribute information may contain missing values. When missing values are encountered, a variety of methods can be used to fill them. For example, if the data distribution is relatively uniform, mean filling can be used, that is, calculating the average of all non-missing values of the attribute and using this average to fill the missing values. If the data distribution is skewed, median filling may be more appropriate, that is, taking the median of all non-missing values of the attribute to fill the missing values. In addition, for some complex data situations, machine learning-based missing value prediction methods can also be used, such as using regression models, decision tree models, etc., to predict missing values based on the values of other attributes.
[0037] Edge relationship information processing: Edge relationship information needs to be checked for abnormal connections. In practical applications, problems such as duplicate connections, self-connections, and incorrectly connected directional relationships may arise. If duplicate connections are found, that is, the same edge is recorded multiple times in the data, these duplicate connections will be deleted and only the single record will be retained. Self-connections, that is, edges connecting a node to itself, are deleted because they do not conform to the actual graph structure semantics in most cases. If there are incorrectly connected directional relationships, such as in a directed graph, where the direction of the edge does not match the actual semantics, these incorrectly connected directional relationships will be corrected to conform to the actual relationship logic.
[0038] Fusion of node attributes and structural information
[0039] The node attribute information and the graph structure information are fused to obtain the fused node feature vector.
[0040] Vectorizing node attribute information: Representing node attribute information as a vector. For example, if a node attribute is numeric, it can be directly converted to the value of the corresponding dimension of the vector. If it is a categorical attribute, it can be converted to a vector using methods such as one-hot encoding. Suppose a node has three attributes: age, gender, and occupation. Age is numeric, gender is categorical (male / female), and occupation is categorical (teacher / doctor / engineer, etc.). For age, its numeric value can be directly used as the value of one dimension of the vector. For gender, if one-hot encoding is used, "male" can be represented as [1,0] and "female" as [0,1]. For occupation, if there are three occupations, "teacher" can be represented as [1,0,0], "doctor" as [0,1,0], and "engineer" as [0,0,1]. These vectors are then concatenated in a certain order to obtain the node attribute vector.
[0041] Graph embedding techniques obtain structural embedding vectors: Graph embedding techniques are used to map nodes in a graph into a low-dimensional vector space to obtain the structural embedding vectors of the nodes. For example, the DeepWalk method generates node sequences by performing random walks on the graph. These node sequences are then treated as sentences in natural language and a model similar to Word2Vec is used to learn the structural embedding vectors of the nodes. The specific steps are as follows: Starting from each node, a random walk of a certain length is performed to generate multiple node sequences. Assuming the random walk length is 5, starting from node A, the possible node sequences generated are ABCDE. These node sequences are then input into the Skip-Gram model (a common Word2Vec model), with the node as the center word and the surrounding nodes as the context words. The model is trained to maximize the co-occurrence probability of the center word and the context words, thereby obtaining the structural embedding vector for each node. Another graph embedding technique, the Node2Vec method, is similar to DeepWalk, but it introduces two parameters, p and q, during the random walk process to control the bias of the walk, enabling it to better capture the local and global structural information of the graph, thereby generating a more representative node structural embedding vector.
[0042] Concatenate the attribute vector and the structural embedding vector: Concatenate the node's attribute vector and structural embedding vector. Assuming the dimension of the node attribute vector is m and the dimension of the structural embedding vector is n, the dimension of the concatenated fused node feature vector is m + n. For example, if the attribute vector is [0.5, 0.3, 0.2] and the structural embedding vector is [0.1, 0.7, 0.2, 0.9], the concatenated fused node feature vector is [0.5, 0.3, 0.2, 0.1, 0.7, 0.2, 0.9].
[0043] Build deep attribute graph models and learn high-level features
[0044] The fused node feature vectors are used to construct a deep attribute graph model, and a graph neural network is used to learn high-order feature representations of nodes.
[0045] Input preparation: The fused node feature vector and the adjacency matrix of the attribute graph are used as the input of the graph neural network. The adjacency matrix is used to describe the connection relationship between nodes in the graph. If there is an edge between node i and node j, the corresponding element A[i][j] in the adjacency matrix is 1 (for undirected graphs, A[j][i] is also 1), otherwise it is 0.
[0046] Graph neural network calculation: The graph neural network in the present invention includes two network layers. During the calculation process of each network layer, the features of adjacent nodes are weighted aggregated and nonlinearly transformed according to the connection relationship between the nodes, and the feature representation of the nodes is updated to learn the high-order feature representation of the nodes.
[0047] Designing a multi-information comparative learning strategy
[0048] A multi-information comparative learning strategy is designed to enhance the deep attribute graph model's ability to recognize node differences by comparing the feature representations of different nodes.
[0049] Construct positive sample pairs and negative sample pairs: Construct positive sample pairs and negative sample pairs. Positive sample pairs are pairs of nodes with close connections and similar attributes in the graph structure. For example, in a social network attribute graph, two user nodes that frequently interact (i.e., the edge connection weight is large) and have similar interests and hobbies (similar attributes) can constitute a positive sample pair. When specifically judging the similarity of node attributes, for numerical attributes, the difference between the two attribute values can be calculated. If the difference is within a certain threshold, the attributes are considered similar; for categorical attributes, if the two attribute values are the same, the attributes are considered similar. Negative sample pairs are pairs of nodes with sparse connections in the graph structure and large attribute differences. Similarly, in the above social network attribute graph, two user nodes that rarely interact (the edge connection weight is small) and have completely different interests and hobbies (large attribute differences) can constitute a negative sample pair.
[0050] Define a contrastive learning loss function: Define a contrastive learning loss function. By minimizing the contrastive learning loss function, the feature representations of positive sample pairs are made closer in feature space, while the feature representations of negative sample pairs are made further apart in feature space. Common contrastive learning loss functions include the InfoNCE loss function.
[0051] Cluster analysis
[0052] Cluster analysis is performed using the learned node feature representation to divide the nodes into different clusters to obtain clustering results.
[0053] Calculate distance or similarity: Calculate the distance or similarity between the learned node feature vectors.
[0054] Nodes are divided into different clusters according to the clustering rules of the specified clustering algorithm. If the K-Means clustering algorithm is used, the number of clusters, K, must first be specified. K nodes are then randomly selected as initial cluster centers. For each node, its distance (or similarity) to the K cluster centers is calculated, and the node is assigned to the cluster with the closest (or most similar) cluster center. Next, the center of each cluster is recalculated, i.e., the mean (or other statistic) of the feature vectors of all nodes in that cluster is used as the new cluster center. The above steps of assigning nodes and updating cluster centers are repeated until the cluster center no longer changes or changes very little, at which point the final clustering result is obtained. If the DBSCAN clustering algorithm is used, two parameters must first be specified: the neighborhood radius ∈ and the minimum number of points, MinPts. For each node, the number of points within its neighborhood with radius ∈ is calculated. If the number of points is greater than or equal to MinPts, the node is considered a core point; if the number of points is less than MinPts but within the neighborhood of a core point, the node is considered a boundary point; otherwise, it is considered a noise point. Starting from a core point, all points in its neighborhood (including core points and boundary points) are divided into a cluster, and then the same operation is performed on other core points that have not been divided. Finally, different clusters and noise points are obtained to complete the cluster analysis.
[0055] Example 1: Social Network User Clustering
[0056] Data acquisition and preprocessing: User data is obtained from a social networking platform and an attribute graph is constructed. Users are nodes, with user attributes such as age, gender, and hobbies. Friendship relationships between users are edges. Missing values in node attributes are filled with the mean, and edge relationships are examined, removing duplicate friendship relationships and user self-following relationships.
[0057] Multi-information fusion: Convert user attributes such as age, gender, and hobbies into vector form. Use one-hot encoding for categorical attributes such as gender and hobbies, and normalize numeric attributes such as age. Use the Node2Vec method to embed the graph structure, setting parameters p = 1, q = 0.5, a random walk length of 8, and the number of walks to 10 to obtain the user's structural embedding vector. Concatenate the attribute vector and the structural embedding vector to obtain the fused user feature vector.
[0058] Deep attribute graph model construction: A deep attribute graph model is constructed using a graph convolutional network (GCN). The fused user feature vector and the social network adjacency matrix are used as input to the GCN, which consists of two network layers. The first network layer maps the input features to 64 dimensions, and the second network layer maps the features to 32 dimensions. In each network layer, based on the friendship relationships between users, the features of adjacent users are weightedly aggregated and nonlinearly transformed (using the ReLU activation function) to update the user's feature representation and learn the user's high-order feature representation.
[0059] Multi-information contrastive learning: constructs positive and negative pairs. Positive pairs are user pairs that are friends and have similar interests (cosine similarity greater than 0.7), while negative pairs are user pairs that are not friends and have significantly different interests (cosine similarity less than 0.3). Define the InfoNCE contrastive learning loss function. By minimizing this loss function, the feature representations of the positive pairs are closer in feature space, while the feature representations of the negative pairs are further apart.
[0060] Cluster analysis: The learned user feature representations were clustered using the K-Means clustering algorithm. The optimal number of clusters, K = 5, was determined using the elbow rule. The distance from each user to the five cluster centers was iteratively calculated, and the user was assigned to the cluster with the closest cluster center. The cluster center was updated until the cluster center remained constant. Ultimately, five user clusters were obtained, corresponding to different interest groups, such as technology enthusiasts, sports enthusiasts, music lovers, gamers, and travel enthusiasts.
[0061] The above is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, can make equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, which should be covered by the scope of protection of the present invention.
Claims
1. A deep attribute graph clustering method based on multi-information comparison, characterized by: The following steps are involved: Acquire social attribute graph data, and clean and standardize node attribute information and edge relationship information in the social attribute graph data; Fuse the node attribute information and the graph structure information to obtain the fused node feature vector; Using the fused node feature vectors to construct a deep attribute graph model, and using a graph neural network to learn high-order feature representations of nodes; Design a multi-information comparative learning strategy to enhance the deep attribute graph model's ability to identify node differences by comparing the feature representations of different nodes; Cluster analysis is performed using the learned node feature representation to divide the nodes into different clusters to obtain clustering results; the clustering results include multiple set interest groups.
2. The deep attribute graph clustering method based on multi-information comparison according to claim 1 is characterized in that: The cleaning and standardization of the node attribute information and edge relationship information in the social attribute graph data is specifically as follows: For node attribute information, if there are missing values, use mean filling, median filling, or missing value prediction methods based on machine learning to fill them; for edge relationship information, check whether there are abnormal connections. If there are abnormal connections, delete duplicate connections and self-connections, and correct the directional relationships of incorrect connections.
3. The deep attribute graph clustering method based on multi-information comparison according to claim 1 is characterized in that: The fusion of node attribute information and graph structure information is as follows: The node attribute information is represented as a vector form, and the graph embedding technology is used to map the nodes in the graph to a low-dimensional vector space to obtain the node's structural embedding vector, and then the node's attribute vector and structural embedding vector are spliced.
4. The deep attribute graph clustering method based on multi-information comparison according to claim 3 is characterized in that: The graph embedding technology is DeepWalk or Node2Vec method.
5. The deep attribute graph clustering method based on multi-information comparison according to claim 1 is characterized in that: The fused node feature vectors are used to construct a deep attribute graph model, and a graph neural network is used to learn the high-order feature representation of the nodes. Specifically: The fused node feature vector and the adjacency matrix of the attribute graph are used as the input of the graph neural network. The graph neural network contains two network layers. During the calculation process of each network layer, the features of adjacent nodes are weighted aggregated and nonlinearly transformed according to the connection relationship between the nodes, and the feature representation of the nodes is updated to learn the high-order feature representation of the nodes; wherein, the graph neural network is a graph convolutional network or a graph attention network.
6. The deep attribute graph clustering method based on multi-information comparison according to claim 1 is characterized in that: The design of the multi-information comparative learning strategy specifically includes: Construct positive sample pairs and negative sample pairs, where the positive sample pairs are node pairs with dense connections and similar attributes in the graph structure, and the negative sample pairs are node pairs with sparse connections and large attribute differences in the graph structure; define a contrastive learning loss function, and by minimizing the contrastive learning loss function, make the feature representations of the positive sample pairs closer in the feature space, and the feature representations of the negative sample pairs farther apart in the feature space.
7. The deep attribute graph clustering method based on multi-information comparison according to claim 1 is characterized in that: Using the learned node feature representation to perform cluster analysis, the nodes are divided into different clusters and the clustering results are as follows: Calculate the distance or similarity between the learned node feature vectors, and divide the nodes into different clusters according to the clustering rules of the set clustering algorithm.
8. The deep attribute graph clustering method based on multi-information comparison according to claim 7 is characterized in that: The clustering algorithm used in the cluster analysis is the K-Means clustering algorithm or the DBSCAN clustering algorithm.