Community division method based on structure enhancement and graph convolutional network
By introducing structural enhancement modules and graph convolutional networks into the deep clustering method, combining the automatic encoder and attention mechanism, the problem of structural information redundancy and insufficient fusion of representation information in the deep clustering method is solved, and a more accurate and robust community discovery effect is achieved.
Patent Information
- Application Number
- CN202510176770.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
The existing deep clustering methods have problems of redundancy in structural information and insufficient fusion of representation information when processing high-dimensional network data, which affects the accuracy and robustness of clustering results.
The community division method based on structure enhancement and graph convolution network is adopted, and the redundant connections in the graph are removed through the structure enhancement module, combined with the automatic encoder and graph convolution neural network, fully utilize the attribute information and structural information of the graph, and fusion of information is carried out through the attention mechanism to optimize the learning process of clustered information.
It improves the accuracy and robustness of community discovery, enhances the quality of graph structure information, and improves the comprehensiveness and reliability of clustering results.
Smart Images

Figure CN120107007A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multi-view clustering and relates to a community division method based on structure enhancement and graph convolutional network. Background Art
[0002] Community detection is an important problem in network analysis, which aims to divide the nodes in the network into multiple communities so that the nodes in the same community have strong connections, while the nodes in different communities have weak connections.
[0003] Community discovery is widely used in social network analysis, bioinformatics, recommender systems and other fields.
[0004] Traditional community detection methods are mainly based on graph theory and statistical inference, such as spectral clustering and modularity optimization.
[0005] However, these methods often have difficulty handling high-dimensional network data and are sensitive to noise and outliers.
[0006] In recent years, deep learning techniques have made remarkable progress in the field of community discovery.
[0007] The deep clustering method uses deep neural networks to learn low-dimensional representations of nodes and performs clustering on this basis. It can effectively process high-dimensional network data and has strong robustness.
[0008] Currently, deep clustering methods mainly use the following two techniques:
[0009] Autoencoder (AE): Autoencoder maps high-dimensional data to a low-dimensional space through an encoder and a decoder and reconstructs the original data to learn the latent features of the data.
[0010] Graph Convolutional Neural Network (GCN): GCN is able to consider the relationship between nodes, learn low-dimensional representations of nodes, and perform clustering based on this.
[0011] However, existing deep clustering methods still have the following problems:
[0012] The structural information in the graph often contains redundancy, such as incorrect connections or repeated edges, which will affect the accuracy of the clustering results.
[0013] Existing methods often process attribute information and structural information separately without fully considering the relationship between the two, resulting in incomplete representation of information.
[0014] In order to solve the above problems, the present invention proposes a community partitioning method based on structure enhancement and graph convolutional network, aiming to improve the accuracy and robustness of community discovery. Summary of the invention
[0015] In view of this, the purpose of the present invention is to provide a community partition method based on structure enhancement and graph convolutional network to solve the problems of structural information redundancy and insufficient representation information fusion existing in deep clustering algorithms.
[0016] In order to achieve the above object, the present invention provides the following technical solutions:
[0017] A community division method based on structure enhancement and graph convolutional network includes the following steps:
[0018] S1: In a social network, each user is a node, and the attribute information of the node includes age, gender and interests; it is represented by feature vectors, and the edges between nodes represent the relationship between users, including friends, attention and comments; by establishing a community partitioning model based on structure enhancement and graph convolutional network, the input includes graph attribute information and graph structure information, that is, the attribute information of users and the relationship between users, and the output is the clustering result based on community structure, that is, the users in the social network are divided into several potential groups;
[0019] S2: Use the structure enhancement module to extract the original graph structure and obtain enhanced structural information;
[0020] S3: Use the autoencoder AE to obtain the potential features of the attribute information of the graph;
[0021] S4: Use the graph convolutional neural network GCN module to integrate the representation learned in the AE module into the GCN module;
[0022] S5: Through the information fusion module, the representation information learned in the AE module and the GCN module is dynamically fused layer by layer;
[0023] S6: Guide and optimize the learning process of clustering information through the self-supervised training module;
[0024] S7: Use the clustering module to obtain the final cluster distribution.
[0025] Furthermore, the structure enhancement module adopts an adversarial learning mechanism to adaptively learn a redundant edge matrix, inputs graph structure information and attribute information into a graph convolutional neural network, and learns an edge deletion matrix through neural network training to ensure the diversity of comparison sample pairs.
[0026] Furthermore, the AE module uses an autoencoder to encode and decode the original image, thereby extracting the attribute information of the image. Through the structure of the autoencoder, the potential features of the image are captured, and the high-dimensional image data is mapped to a low-dimensional space, thereby improving the expression ability of the image attribute information.
[0027] Furthermore, the GCN module uses a graph convolutional neural network to concentrate all representations learned in the AE module into the GCN, considering the relationship between the data itself and the samples, and providing embedded representations for downstream tasks.
[0028] Furthermore, the information fusion module adopts an attention mechanism fusion strategy to input the acquired attribute representation information and enhanced structural information into a fully connected neural network. Through the training of the neural network, the fusion weights between different information sources are learned, and by dynamically adjusting the contribution of each information source, the model can adaptively optimize the information fusion strategy in different tasks.
[0029] Furthermore, a supervised loss method is designed in the self-supervised training module to guide and optimize the learning process of clustering information. The loss function constrains the clustering results output by the model, promotes the model to gradually improve the representation ability between modules during the training process, and ensures that the features of different modules can work together.
[0030] Furthermore, the clustering module uses a layer of perceptron to directly obtain the cluster distribution of the representation, and avoids suboptimal clustering performance by combining the optimization process of representation learning and node clustering.
[0031] The beneficial effects of the present invention are:
[0032] (1) Through the structure enhancement module, the adversarial learning mechanism is used to learn the redundant edge matrix, remove the erroneous or redundant connections in the graph, enhance the structural information of the graph, and avoid the influence of redundant information on the clustering results.
[0033] (2) Combining autoencoders and graph convolutional neural networks, we can make full use of the attribute information and structural information of the graph to obtain a more comprehensive and richer node representation and improve the accuracy of the clustering results.
[0034] (3) The attention mechanism is used for dynamic information fusion, and the fusion weights are adaptively adjusted according to the importance of different features, so that the model can better adapt to different tasks and improve the robustness of the clustering results.
[0035] (4) Design a dual self-supervision mechanism to optimize data representation by learning high-confidence assignments, improve clustering effects and module representation capabilities, and prevent the model from falling into a local optimal solution.
[0036] (5) Through the one-step clustering module, the cluster distribution of the representation is directly obtained, avoiding the suboptimal clustering performance caused by the separation of representation learning and node clustering optimization processes, and improving the reliability of clustering results.
[0037] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0039] Figure 1 Model diagram of the community partitioning method based on structure enhancement and graph convolutional network. DETAILED DESCRIPTION
[0040] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0041] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0042] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0043] The community division method of the present invention under multi-view fusion information, Figure 1 It is a community division method model based on structural enhancement and graph convolutional network. The whole model consists of 6 modules:
[0044] The structure enhancement module uses an adversarial learning mechanism to adaptively learn the redundant edge matrix, which inputs the graph structure information and attribute information into the graph convolutional neural network. Through the training of the neural network, the module can learn the edge deletion matrix to ensure the diversity of the comparison sample pairs.
[0045] The AE module uses an autoencoder (AE) to encode and decode the original image to extract the attribute information of the image. Through the structure of the autoencoder, the potential features of the image can be effectively captured, and the high-dimensional image data can be mapped to a low-dimensional space, thereby improving the expression ability of the image attribute information.
[0046] The GCN module uses a graph convolutional neural network to concentrate all representations learned in the AE module into the GCN, which not only considers the data itself, but also the relationship between samples, providing a better embedding representation for downstream tasks.
[0047] The information fusion module adopts the attention mechanism fusion strategy to input the acquired attribute representation information and enhanced structural information into the fully connected neural network. Through the training of the neural network, the module can learn the fusion weights between different information sources, and by dynamically adjusting the contribution of each information source, the model can adaptively optimize the information fusion strategy in different tasks. Ultimately, this fusion method helps to improve the accuracy and robustness of community discovery, thereby obtaining a clustering representation that is more suitable for community discovery.
[0048] The self-supervised training module designs a supervised loss method to more effectively guide and optimize the learning process of clustering information. This loss function can promote the model to gradually improve the representation ability between modules during the training process by constraining the clustering results output by the model, ensuring that the features of different modules can work together, thereby improving the overall performance.
[0049] The clustering module uses a layer of perceptron to directly obtain the cluster distribution of representations, and avoids suboptimal clustering performance through the optimization process of joint representation learning and node clustering.
[0050] The specific implementation steps of the present invention are as follows:
[0051] Step 1: Given an attribute graph G = (V, E, X), where is a vertex set with n vertices, E is an edge set, X = [x 1 ,x 2 ,…,x n ] T is the feature matrix. The topological structure of graph G can be represented by the adjacency matrix It means that if v ij =1, indicating vertex vi To the vertex v j There is an edge between them. represents the degree matrix of A, where Represents vertex v i The degree.
[0052] Step 2: Use the structure enhancement module to extract the enhanced structural information representation of the original graph structure A The specific steps include:
[0053] Step 21: In order to remove the wrong or redundant connections in the original graph, a structure enhancement module is designed to learn the clustering-oriented structure graph. A GCN-Encoder encoder is used to receive the normalized adjacency matrix and node attribute matrix X as input, and output graph embedding Z∈R N×d’ The calculation formula is as follows:
[0054]
[0055] in is the original feature matrix, is the normalized adjacency matrix, and the calculation formula is as follows:
[0056]
[0057] Step 22: Then use the graph embedding Z to generate the edge embedding matrix Where E m =C(z i ,z j ), E m = C(.,.) is a concatenation operation, and then the obtained edge embedding E is input into a 1-layer MLP with a sigmoid activation function to obtain the edge-oriented weight vector As shown below:
[0058] w i =sigmoid(MLP(E i ))
[0059] Among them, w i Represents the original edge e i probability of retention.
[0060] Step 23: Next, Converted into edge-oriented weight matrix W'∈R N×N , thus obtaining the enhanced adjacency matrix The calculation formula is as follows:
[0061]
[0062] Step 3: Use the AE autoencoder to encode and decode the original feature matrix X to obtain the attribute information of the graph. Specifically, the following steps are included:
[0063] Step 31: Input the feature matrix X of the original image into the encoder of AE. The calculation formula is as follows:
[0064]
[0065] Among them, H (0) =X is the original feature matrix, is the parameter matrix learned by the encoder layer l. It is an activation function of a fully connected layer similar to Relu or Sigmoid function.
[0066] Step 32: Pass the output of the encoder to the decoder. The calculation formula is as follows:
[0067]
[0068] in is the learnable parameter matrix of the decoder layer l, It is an activation function similar to RELU, and the output of the encoder is the reconstructed feature matrix
[0069] Step 33: Calculate the reconstruction loss of the AE module, which is used to measure the difference between the input feature matrix and the reconstructed feature matrix. The specific calculation formula is as follows:
[0070]
[0071] Where X is the original feature matrix, is the feature matrix reconstructed by the decoder. The loss function guides model learning by minimizing this difference, ensuring that the encoder can extract effective graph attribute information.
[0072] Step 4: Use the GCN module to integrate the representation learned in the AE module into the GCN module. Then the representation that can be learned by GCN will be able to accommodate two different types of information, namely the data itself and the relationship between the data. The specific steps include:
[0073] Step 41: Convert the feature matrix X of the original graph and the normalized adjacency matrix and the enhanced structural matrix Input into the GCN encoder, the calculation formula is as follows:
[0074]
[0075] Among them, Z( 0 )=X is the original feature matrix.
[0076] Step 5: The representation information learned in the AE module and the CGN module is dynamically fused layer by layer through the attention mechanism so that the model can comprehensively consider the influence of different features and further improve the community discovery effect of the graph. The specific steps include:
[0077] Step 51: First, the attribute information representation Z and the enhanced neighborhood structure representation information H are concatenated together to form a connected feature vector. Then, the relationship between the connected features is captured through the fully connected layer, and then mapped through the nonlinear activation function LeakyReLu. Finally, the concatenated vector is normalized using softmax and L2 regularization to obtain the attention coefficient matrix C. The formula is as follows:
[0078]
[0079] where C∈R N×d Attention coefficient matrix, C1, C2 represent the fused attention coefficient vectors of H and Z respectively, and W is the learned parameter matrix.
[0080] Step 52: Next, H and Z are fused through the Hadamard product to obtain the final feature representation F. The formula is as follows:
[0081] Z'=C⊙H+(1 N×d -C)⊙Z
[0082] Among them, ⊙ represents the Hadamard inner product operation. The final fused feature representation Z' contains important information from the AE module and the GCN module. By introducing this attention mechanism, the obtained fused representation can more effectively capture the key information from the two modules, thereby improving the performance of clustering and community discovery. This fused representation provides more favorable feature support for subsequent tasks.
[0083] Step 6: In order to improve the clustering effect and the representation ability of the module, a dual self-supervision mechanism is designed so that the model can learn a more suitable clustering distribution. The specific steps include:
[0084] Step 61: Use Student's t distribution as the kernel to measure the similarity between the data representation hi and the cluster center vector μj. The specific calculation formula is as follows:
[0085]
[0086] Among them, h i Represents H( l ) i-th row, μ j is the jth cluster center initialized by k-means, v is the degrees of freedom of the Student's t distribution, qij represents the probability that node i belongs to cluster j, that is, soft assignment Q = [q ij ].
[0087] Step 62: After obtaining the clustering result distribution Q, the goal is to optimize the data representation by learning high confidence assignments. Specifically, it is hoped that the data representation is closer to the cluster center, thereby improving the cohesion of the cluster. Therefore, the target distribution P is calculated as follows:
[0088]
[0089] where f j =∑ i q ij is the soft cluster frequency.
[0090] Step 63: For training the GCN module, one possible approach is to use the cluster assignment as the true value label [3]. However, this strategy will lead to noisy and trivial solutions and cause the collapse of the entire model. As mentioned earlier, the GCN module also provides a cluster assignment distribution Z. Therefore, the distribution Z can be supervised by the distribution P:
[0091]
[0092] Step 7: Many existing clustering methods represent that the optimization process of learning and node clustering is separated, which leads to suboptimal clustering performance. To solve this problem, a one-step clustering module is designed. It includes the following steps:
[0093] Step 71: After obtaining the graph embeddings of the two views, they are converted into a k-dimensional clustering space by inputting Z1 and Z2 into a 1-layer multi-layer perceptron MLP with a softmax activation function, where K represents the number of clusters. The specific calculation formula is as follows:
[0094] C v =softmax(MLP(Z v )),v∈{1,2}
[0095] Among them, C v ∈R N×K The clustering indicator matrix in the vth view. 1 and C 2 The clustering result is obtained by taking the average value of .
[0096] Step 72: To alleviate the problem of the encoder capturing redundant information when estimating the consistency of sample pairs, a redundancy reduction strategy is introduced in the latent space by forcing the cross-view correlation matrix to approximate the identity matrix:
[0097]
[0098] This term reduces the redundancy of embeddings across dual views in the corresponding graph. This approach can minimize the redundant information in the embeddings and retain more discriminative features. Therefore, it makes the learned representation less affected by irrelevant information, thereby ensuring the quality of the latent space for subsequent clustering tasks.
[0099] Step 73: The final optimization objective loss function is:
[0100] L total =L ae +L gcn +L mse
[0101] Example
[0102] (1) Data preparation: We use the public social network dataset Cora as the input attribute graph G = (V, E, X), where the number of nodes n = 2708 and the adjacency matrix A ∈ R 2708×2708 , feature matrix X∈R 2708×1433 .
[0103] (2) Structural enhancement:
[0104] Normalized adjacency matrix
[0105] Generate graph embedding Z through GCN encoder, calculate edge weight W', and get enhanced adjacency matrix
[0106] (3) Attribute extraction:
[0107] The AE encoder maps XX to a low-dimensional space and reconstructs the loss
[0108] (4) Information Fusion:
[0109] Concatenate the AE output H and the GCN output Z, and perform weighted fusion through the attention coefficient C to obtain Z′.
[0110] (5) Self-supervised training:
[0111] Calculate the student t distribution Q and target distribution P, and optimize the KL loss L gcn ;
[0112] (6) Clustering output:
[0113] Dual View Embedding Z 1 and Z 2 After MLP mapping, the average is taken and the final clustering result is output.
[0114] (7) Performance evaluation:
[0115] On the Cora dataset, the clustering accuracy ACC reaches 86.3%, an increase of 12.5% over the baseline model.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A community partitioning method based on structure enhancement and graph convolutional network, characterized by: The following steps are involved: S1: In a social network, each user is a node, and the node's attribute information includes age, gender, and interests; The edges between nodes represent the relationships between users, including friends, followers, and comments, through feature vector representation. A community partitioning model based on structure enhancement and graph convolutional network is established. The input includes graph attribute information and graph structure information, i.e., user attribute information and the relationships between users. The output is the clustering result based on community structure, i.e., the users in the social network are divided into several potential groups. S2: Use the structure enhancement module to extract the original graph structure and obtain enhanced structural information; S3: Use the autoencoder AE to obtain the potential features of the attribute information of the graph; S4: Use the graph convolutional neural network GCN module to integrate the representation learned in the AE module into the GCN module; S5: Through the information fusion module, the representation information learned in the AE module and the GCN module is dynamically fused layer by layer; S6: Guide and optimize the learning process of clustering information through the self-supervised training module; S7: Use the clustering module to obtain the final cluster distribution.
2. The community partitioning method based on structure enhancement and graph convolutional network according to claim 1, characterized in that: The structure enhancement module adopts an adversarial learning mechanism to adaptively learn the redundant edge matrix, inputs the graph structure information and attribute information into the graph convolutional neural network, and learns the edge deletion matrix through the training of the neural network to ensure the diversity of the comparison sample pairs.
3. The community partitioning method based on structure enhancement and graph convolutional network according to claim 1, characterized in that: The AE module uses an autoencoder to encode and decode the original image, thereby extracting the attribute information of the image. Through the structure of the autoencoder, the potential features of the image are captured, and the high-dimensional image data is mapped to a low-dimensional space, thereby improving the expression ability of the image attribute information.
4. The community partitioning method based on structure enhancement and graph convolutional network according to claim 1, characterized in that: The GCN module uses a graph convolutional neural network to concentrate all representations learned in the AE module into the GCN, considering the relationship between the data itself and the samples, and providing embedded representations for downstream tasks.
5. The community partitioning method based on structure enhancement and graph convolutional network according to claim 1, characterized in that: The information fusion module adopts an attention mechanism fusion strategy to input the acquired attribute representation information and enhanced structural information into a fully connected neural network. Through the training of the neural network, the fusion weights between different information sources are learned. By dynamically adjusting the contribution of each information source, the model can adaptively optimize the information fusion strategy in different tasks.
6. The community partitioning method based on structure enhancement and graph convolutional network according to claim 1, characterized in that: A supervised loss method is designed in the self-supervised training module to guide and optimize the learning process of clustering information. The loss function constrains the clustering results output by the model, promotes the model to gradually improve the representation ability between modules during the training process, and ensures that the features of different modules can work together.
7. The community partitioning method based on structure enhancement and graph convolutional network according to claim 1, characterized in that: The clustering module uses a layer of perceptron to directly obtain the cluster distribution of representations, and avoids suboptimal clustering performance through the optimization process of joint representation learning and node clustering.
Citation Information
Cited By
Social network user community identification method based on multi-scale graph comparative learning
CN120318003A