Network community division method based on dynamic adjustment and neighbor comparison module
By introducing dynamic adjustment and neighbor comparison modules in graph comparison learning, the problem of existing methods ignoring the limitations of node local structure and encoder output characteristics is solved, and higher clustering accuracy and model generalization capabilities are achieved.
Patent Information
- Application Number
- CN202411760607.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-03
AI Technical Summary
The existing graph comparison learning method ignores the local structural homogeneity of nodes in graph data, is limited to the encoder output characteristics, and comparison learning on large-scale graphs requires a large number of negative samples, resulting in high computational cost and poor generalization ability.
A method of network community division based on dynamic adjustment and neighbor comparison module is proposed. The enhanced graph is generated through the graph enhancement module, the features are extracted using a dual-layer encoder, the adaptive adjustment module dynamically adjusts the graph structure, and the loss contribution is evaluated through the neighbor comparison module, and the model is finally optimized through the total loss function.
It improves the accuracy of the model in clustering and prediction, enhances the scalability and generalization capabilities of the model, can better capture the local and global features of the graph, and reduces noise interference.
Smart Images

Figure CN119693171B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to network social group division technology, and in particular to a network community division method based on dynamic adjustment and neighbor comparison modules. Background Art
[0002] In recent years, with the widespread success of self-supervised learning in the fields of computer vision (CV) and natural language processing (NLP), contrastive learning has gradually become a mainstream self-supervised learning method. The core idea of contrastive learning is to distinguish between similar and dissimilar sample pairs, so that the model can automatically learn the intrinsic feature representation of the sample, and then perform well in downstream tasks. Inspired by this success, researchers began to try to expand contrastive learning to the field of graph data, forming Graph Contrastive Learning (GCL).
[0003] GCL, its core technology, combines graph neural networks (GNNs) and contrastive learning. It maximizes the similarity of the same node in different views and minimizes the similarity of different nodes by constructing graph contrast losses for different views. Most current technologies focus on graph data enhancement, encoder extraction representation, calculation of contrastive losses, and finally optimization of models. However, existing similar technologies have several key defects:
[0004] 1. Ignoring the homogeneity of the local structure of nodes in graph data
[0005] Existing GCL methods usually ignore the local structural homogeneity characteristics of nodes in graph data. For example, in InfoNCE loss, adjacent nodes are often regarded as negative samples, which is contrary to the assumption of graph homogeneity. In actual homogeneous graphs, adjacent nodes often share the same semantic labels, which means that adjacent nodes have potential similarities and should be regarded as positive sample pairs rather than negative sample pairs.
[0006] 2. Limited to encoder output characteristics:
[0007] Most existing GCL methods usually only focus on the output features of the graph encoder and ignore the global topological structure of the graph. This approach results in the model failing to fully utilize the structural information of the graph data.
[0008] 3. Large-scale comparative learning on graphs:
[0009] Existing graph contrastive learning often requires a large number of negative samples to learn node representations. However, in actual scenarios, the scale of graphs is often very large. Therefore, a large number of negative samples requires huge memory and computational costs, which limits the generalization ability of the model. Summary of the invention
[0010] Purpose of the invention: The purpose of the present invention is to solve the deficiencies in the prior art and to provide a network community division method based on dynamic adjustment and neighbor comparison modules.
[0011] Technical solution: A network community division method based on dynamic adjustment and neighbor comparison modules of the present invention comprises the following steps:
[0012] Step 1: Obtain social network data and construct the original graph G = (X, A), where A represents the adjacency matrix of the social network, representing the connection relationship between each node, X represents the feature matrix of the nodes in the social network, and the nodes represent the users in the social network;
[0013] The graph enhancement module divides the original graph G = (X, A) into communities based on the sample learning positive and negative example method, and generates two new enhanced graphs G 1 =(X1, A 1 ) and G 2 =(X 2 , A 2 ), the specific method is: randomly delete the unimportant edges in the original graph to get A 1 and A 2 , randomly delete unimportant attributes in the original graph to get X 1 and X 2 ;
[0014] Step 2: Enhance the graph G 1 =(X 1 , A 1 ) and the enhanced graph G 2 =(X 2 , A 2 ) inputs the shared parameters of the two-layer encoder f θ , through the encoder f θ Perform feature extraction on the two enhanced graphs respectively to obtain the corresponding node features Z 1 and Z 2 ;
[0015] Step 3: The feature Z obtained in step 2 1 and Z 2 Input adaptive adjustment module, feature Z 1 and Z 2 Dynamically adjust and reconstruct respectively to obtain a new adjacency matrix A′ 1 and A′ 2 ; A′ 1 and A 1 Compare and 2 and A′ 2 By comparison, we get the corresponding cross entropy loss k 1 and k 2 ; The expression is as follows:
[0016]
[0017]
[0018] k 1 =[-log(A′ 1pos )-log(1-A′ 1ne )]
[0019] k 2 =[-log(A′ 2po )-log(1-A′ 2neg )]
[0020] In the above formula, T refers to the matrix transpose, σ is the Sigmoid function, and A′ 1pos Refers to A' 1 , A′ 1neg A′ is generated by negative sampling 1po The edge that does not exist in the view; node i represents the anchor point, and node j refers to the first-order neighbor node of node i in the view; It refers to the enhanced graph G 1 The i-node features, It refers to the enhanced graph G 1 The node features of node j connected to node i (node i and node j are connected by an edge); It refers to the enhanced graph G 2 The i-node features, It refers to the enhanced graph G 2 The node features of node j connected to node i (node i and node j are connected by an edge;
[0021] Step 4: Use the neighbor comparison module to evaluate the contribution of each neighbor to the loss calculation and calculate the comparison loss L N When , the similarity of the adjacent nodes of the anchor point is introduced and added to the positive sample pair; the positive nodes include the same nodes of the anchor points in different views and the neighbor nodes of the anchors within the view and across different views; the negative nodes include the non-neighbors of the anchors within the view and across different views; L N The calculation formula is as follows:
[0022]
[0023] In the above formula, It refers to the similarity of the adjacent nodes of the anchor point;
[0024] Finally, the total loss function is used to evaluate the difference between the model's predicted value and the true value; the total loss includes the contrast loss L N , reconstruction loss k 1 and reconstruction loss k 2 ;
[0025] Contrastive loss L N :
[0026]
[0027] In the above formula, It refers to the similarity between the same nodes in different views. It refers to the similarity between the anchor point and its neighbor nodes in the same view. It refers to the similarity between the anchor point and its neighbor nodes in different views; the reconstruction loss k 1 and reconstruction loss k 2 :
[0028] k 1 =[-log(A′ 1pos )-log(1-A′ 1neg )]
[0029] k 2 =[-log(A′ 2pos )-log(1-A′ 2neg )]
[0030] At the end of training, the graph structure is dynamically adjusted according to the node features to help the encoder better train the model; by reconstructing two approximate original adjacency matrices, the reconstruction loss is calculated to minimize the difference between the reconstructed adjacency matrix and the original adjacency matrix.
[0031] Furthermore, the detailed method of step 1 is:
[0032] Step 1.1: construct an original graph based on social network data, and divide the original graph G = (X, A) into different communities according to the community detection algorithm;
[0033] Step 1.2: For the adjacency matrix A, first delete the edges between communities, then delete the edges of communities with weak community strength S, and get the new adjacency matrix A. 1 and A 2 ; The method is
[0034] Use the scoring function W(e) to calculate each edge e∈ε in the community c c The weight w e ;
[0035]
[0036]
[0037] H∈{0,1} n×c is the community indicator matrix, H i,c =1[v i ∈c] represents which community the i-th node belongs to;
[0038]
[0039] The edge mask me is obtained by sampling independently from two Bernoulli distributions. The edge mask me (using this mask to filter the edges and only retain the randomly selected edges, and the edge deletion is completed) can determine which edges are discarded. and are hyperparameters, and both have values between 0 and 1, which will eventually result in the adjacency matrix A 1 and A 2 ;
[0040]
[0041] Step 1.3: For the feature matrix X, remove the attributes with less influence to obtain a new feature matrix X 1 and X 2 ; The method is:
[0042] First, define a W a Represents the probability of removing the attribute of each node in the social network:
[0043]
[0044] S refers to the community strength, and abs is the absolute value to ensure it is non-negative; It is a one-dimensional normalization operation;
[0045] Then, by W a Compute the attribute mask m a :
[0046]
[0047] in and is a hyperparameter, ranging from 0 to 1, unequal, and then determines two different attribute matrices X 1 and X 2 ; It is the same as edge deletion:
[0048] and
[0049] Step 1.4: After the changes in edges and attributes, the original graph G = (X, A) generates two new enhanced graphs G 1 =(X 1 , A 1 ) and G 2 =(X 2 , A 2 ).
[0050] In step 1.3
[0051] W a For the probability of each attribute being removed, then W a Large attributes will not be removed, but small ones will be removed. The judgment is based on the following formula:
[0052]
[0053] The above Bernoulli distribution is calculated and It is true or false, and then delete the attribute based on these two, and delete the attribute based on the final true and false.
[0054] Furthermore, the community strength S c The calculation formula is as follows:
[0055]
[0056] Among them, ε is the sum of all the edges in the original graph, ε c is the sum of all edges in the community, v is the node in the original graph, and d(v) is the sum of the degrees of the nodes. The definition of strong and weak communities is to compare S c size.
[0057]
[0058] The edge weights within a community are positive numbers, while the edge weights between communities are negative numbers. Because of the negative sign, the size is obvious.
[0059] Furthermore, the encoder f θ Based on a two-layer graph convolutional neural network GCN, sharing the same parameters;
[0060] Encoder f θ The structure is as follows:
[0061]
[0062]
[0063]
[0064] Where Z (1) It represents the feature vector of the node at layer l, W (1) represents the parameters of the first layer of convolution, represents the adjacency matrix of the feature matrix, represents the degree matrix of the adjacency matrix, I represents the identity matrix, σ represents the activation function, represents the Laplacian matrix;
[0065] Finally, the feature representation of the node is as follows;
[0066] like Indicates that there is an edge between node i and node j;
[0067] like Indicates that there is no edge between node i and node j.
[0068] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0069] 1. Accuracy
[0070] Traditional graph contrast learning methods usually mistake adjacent nodes as negative samples. However, the present invention introduces a neighbor contrast module and weights the nodes according to their similarity scores, so that the model can more accurately distinguish between positive and negative sample pairs, thereby improving the clustering and prediction accuracy of the model.
[0071] 2. Scalability
[0072] Through the dynamic adjustment mechanism of the adaptive adjustment module, the present invention effectively learns in different types of graph structures (such as social networks, knowledge graphs, and biological networks) without complex preprocessing of the data. This design enables the model to quickly adapt and maintain good performance when facing various graph data sets.
[0073] 3. Improve the generalization ability of the model
[0074] The present invention effectively suppresses noise interference in the original image by better capturing the local and global features of the image, and reduces the risk of overfitting the model to abnormal samples. This not only improves the performance of the model on the training set, but also significantly improves the generalization ability on the test set. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 It is a schematic diagram of the overall process framework of the present invention;
[0076] Figure 2 It is a schematic diagram of the adaptive adjustment module in the present invention;
[0077] Figure 3 It is a schematic diagram of a neighbor comparison module in the present invention;
[0078] Figure 4 It is a schematic diagram of the effect in the embodiment. DETAILED DESCRIPTION
[0079] The technical solution of the present invention is described in detail below, but the protection scope of the present invention is not limited to the embodiments.
[0080] In existing social network computer technology, traditional community partitioning algorithms often have deficiencies in feature extraction, especially in capturing fine-grained characteristics and relationships between nodes in complex networks, which leads to the problem of node misjudgment in community partitioning. For example, some nodes in community 1 may be misclassified into community 2 due to feature ambiguity, which reduces the accuracy of the overall partitioning and its practical application value.
[0081] To solve the above technical defects, the present invention captures the interaction information between nodes and neighbors, dynamically optimizes the mechanism, adjusts the feature vectors of nodes in the community, eliminates noise and redundant information to the greatest extent, and uses the optimized node features to divide the community through the k-means algorithm. The present invention extracts more delicate features of nodes in the community, thereby helping to better divide the community. For example, on e-commerce or social platforms, by accurately dividing user communities, more personalized recommendation services can be provided, such as recommending related products or content to users in the same interest community.
[0082] like Figure 1 As shown, a network community division method based on dynamic adjustment and neighbor comparison module of the present invention comprises the following steps:
[0083] Step 1: Obtain social network data and construct the original graph G = (X, A), where A represents the adjacency matrix of the social network, representing the connection relationship between each node, X represents the feature matrix of the nodes in the social network, and the nodes represent the users in the social network;
[0084] The graph enhancement module divides the original graph G = (X, A) into communities based on the sample learning positive and negative example method, and generates two new enhanced graphs G 1 =(X 1 ,A 1 ) and G 2 =(X 2 ,A 2 ), the specific method is: randomly delete the unimportant edges in the original graph to get A 1 and A 2 , randomly delete unimportant attributes in the original graph to get X 1 and X 2 ;
[0085] Step 2: Enhance the graph G 1 =(X 1 , A 1 ) and the enhanced graph G 2 =(X 2 , A 2 ) inputs the shared parameters of the two-layer encoder f θ , through the encoder f θ Perform feature extraction on the two enhanced graphs respectively to obtain the corresponding node features Z1 and Z 2 ;
[0086] Step 3: The feature Z obtained in step 2 1 and Z 2 Input adaptive adjustment module, feature Z 1 and Z 2 Dynamically adjust and reconstruct respectively to obtain a new adjacency matrix A′ 1 and A′ 2 ; A′ 1 and A 1 Compare and 2 and A′ 2 By comparison, we get the corresponding cross entropy loss k 1 and k 2 ; The expression is as follows:
[0087]
[0088]
[0089] k 1 =[-log(A′ 1pos )-log(1-A′ 1neg )]
[0090] k 2 =[-log(A′ 2pos )-log(1-A′ 2ne )]
[0091] In the above formula, T refers to the matrix transpose, σ is the Sigmoid function, and A′ 1pos Refers to A' 1 , A′ 1neg A′ is generated by negative sampling 1pos The edge that does not exist in the view; node i represents the anchor point, and node j refers to the first-order neighbor node of node i in the view; It refers to the enhanced graph G 1 The i-node features, It refers to the enhanced graph G 1 The node characteristics of node j connected to node i; It refers to the enhanced graph G 2 The i-node features, It refers to the enhanced graph G 2 The node characteristics of node j connected to node i in , such as Figure 2 As shown;
[0092] Step 4: Figure 3 As shown, the neighbor comparison module is used to evaluate the contribution of each neighbor to the loss calculation and calculate the contrast loss L NWhen , the similarity of the adjacent nodes of the anchor point is introduced and added to the positive sample pair; the positive nodes include the same nodes of the anchor points in different views and the neighbor nodes of the anchors within the view and across different views; the negative nodes include the non-neighbors of the anchors within the view and across different views; L N The calculation formula is as follows:
[0093]
[0094] In the above formula, It refers to the similarity of the adjacent nodes of the anchor point;
[0095] Finally, the total loss function is used to evaluate the difference between the model's predicted value and the true value; the total loss includes the contrast loss L N , reconstruction loss k 1 and reconstruction loss k 2 ; Contrastive loss L N :
[0096]
[0097] In the above formula, It refers to the similarity between the same nodes in different views. It refers to the similarity between the anchor point and its neighbor nodes in the same view. It refers to the similarity between the anchor point and its neighbor nodes in different views; the reconstruction loss k 1 and reconstruction loss k 2 :
[0098] k 1 =[-log(A′ 1pos )-log(1-A′ 1neg )]
[0099] k 2 =[-log(A′ 2pos )-log(1-A′ 2neg )]
[0100] By reconstructing two approximate original adjacency matrices, the reconstruction loss is calculated to minimize the difference between the reconstructed adjacency matrix and the original adjacency matrix.
[0101] The detailed method of step 1 of this embodiment is:
[0102] Step 1.1: construct an original graph based on social network data, and divide the original graph G = (X, A) into different communities according to the community detection algorithm;
[0103] Step 1.2: For the adjacency matrix A, first delete the edges between communities, then delete the edges with weak community strength S, and get the new adjacency matrix A. 1 and A 2 ; The method is
[0104] Use the scoring function W(e) to calculate each edge e∈ε in the community c c The weight w e ;
[0105]
[0106]
[0107] H∈{0,1} n×c is the community indicator matrix, H i,c =1[v i ∈c] represents which community the i-th node belongs to;
[0108]
[0109] The edge mask me is obtained by independent sampling from two Bernoulli distributions, and the edge masks me are used to determine which edges are discarded. and is a hyperparameter, which will eventually lead to the final adjacency matrix A 1 and A 2 ;
[0110]
[0111] Step 1.3: For the feature matrix X, remove the attributes with less influence to obtain a new feature matrix X 1 and X 2 ; The method is:
[0112] First, define a W a Represents the probability of removing the attribute of each node in the social network:
[0113]
[0114] , abs is the absolute value to ensure non-negative; It is a one-dimensional normalization operation;
[0115] Then, by W a Compute the attribute mask m a :
[0116]
[0117] in and is a hyperparameter, ranging from 0 to 1, unequal, and then determines two different attribute matrices X 1 and X 2 ; It is the same as edge deletion:
[0118] and
[0119] Step 1.4: After the changes in edges and attributes, the original graph G = (X, A) generates two new enhanced graphs G 1 =(X 1 , A 1 ) and G 2 =(X 2 , A 2 ).
[0120] In this example, the community strength S c The calculation formula is as follows:
[0121]
[0122] Among them, ε is the sum of all the edges in the original graph, ε c is the sum of all edges in the community, v is the node in the original graph, and d(v) is the sum of the degrees of the nodes. The definition of strong and weak communities is to compare S c size.
[0123] The encoder f in this embodiment θ Based on a two-layer graph convolutional neural network GCN, sharing the same parameters;
[0124] Encoder f θ The structure is as follows:
[0125]
[0126]
[0127]
[0128] Among them, H (1) It represents the feature vector of the node at layer l, W (1) represents the parameters of the first layer of convolution, represents the adjacency matrix of the feature matrix, represents the degree matrix of the adjacency matrix, I represents the identity matrix, σ represents the activation function, represents the Laplacian matrix;
[0129] Finally, the feature representation of the node is as follows;
[0130] like Indicates that there is an edge between node i and node j;
[0131] like Indicates that there is no edge between node i and node j.
[0132] like Figure 2As shown, the original graph is transformed into two enhanced graphs through the graph enhancement module, and then the two enhanced graphs are passed through the encoder to obtain node features, and a new graph is reconstructed based on the node features. The adaptive adjustment module of the present invention has different features for each node in each training process, and the reconstructed graph is also different.
[0133] This embodiment uses a neighbor comparison module to evaluate the contribution of each neighbor to the loss calculation (the anchor node can pay attention to the neighbor node information. The higher the neighbor node feature information is, the greater the impact on the graph representation learning is, and thus the greater the contribution to the loss calculation is. Because the information of the neighbor node is a positive sample, it has an enhancement effect)
[0134] To further verify the technical effect of the present invention, as shown in Table 1, the accuracy results of the node clustering task of the technical solution of the present invention and the existing self-supervised method on four benchmark data sets are compared.
[0135] Table 1
[0136]
[0137] It can be seen from Table 1 that the models of the present invention surpass other models, have good community clustering results, and significantly improve the performance of graph contrast learning.
[0138] The technical solution of the present invention can be used to discover closely related user groups in a social network, where these user groups have common interests, geographical locations or behavior patterns.
[0139] By analyzing the interactive relationships between users, such as likes, comments and follows, community clustering can dig out user interest groups; for example, on a music platform, users who like the same type of music may be clustered into a community. The community structure obtained by the division of the present invention helps to identify key nodes of information dissemination (such as opinion leaders), thereby optimizing advertising or public opinion dissemination strategies; at the same time, it can recommend potential social relationships with similar interests or common friends to users, thereby improving user experience.
[0140] This embodiment divides the members of a karate club, wherein four different colors represent four communities, wherein the community strength of the blue and green communities is large, the community strength of the yellow and red communities is relatively weak, and the edges of the blue and green communities are denser and more compact.
[0141] In the original graph G = (X, A), X represents the attribute characteristics of each person (the dimension is 34×34), A represents the relationship between each person (the dimension is 34×34), there are 34 people in the club, and each person has 34 characteristics, and there are 156 edges in the graph.
[0142] Next, the technical solution of the present invention is used to divide the original graph G = (X, A) into communities. The specific method is as follows:
[0143] Step 1: Calculate the community strength of the four communities according to the formula, and get two new views G 1 =(X 1 , A 1 ) and G 2 =(X 2 , A 2 )
[0144] Step 2: Feed the two new enhanced graphs into the shared encoder GCN to obtain the feature representation of the node, and then optimize the node representation through the dynamic adjustment module and the neighbor comparison module.
[0145] Step 3, then use the K-means algorithm to transform Z 1 and Z 2 As input, minimize the distance of each node to the cluster center:
[0146]
[0147] By adjusting each point cluster to assign c i and update the cluster center u k , minimize the sum of squared distances from points to cluster centers, with the goal of making points in the same cluster as close to the center as possible, thus forming a compact clustering structure. Compare the differences between C1 and C2 to evaluate model performance.
[0148] In the above formula, N refers to the total number of nodes. In this embodiment, there are 34 users in total, so the value is 34; K refers to the number of clusters. Assuming that this embodiment divides the social network according to the interests of users, and there are 10 hobbies, then the number of clusters is 10. The final result of this embodiment is to divide the 34 users into 10 hobbies. i Indicates node characteristics; u k represents the center of cluster k, that is, the centroid of all points in the cluster; c i Indicates the cluster label to which node i belongs, and its value range is {1, 2, ...K}; 1[c i =k] is the indicator function, if c i = k, then the value is 1, otherwise it is 0; || z i -u k || 2 Represents point z i To cluster center u k The final division effect of this embodiment is as follows: Figure 4 shown.
Claims
1. A network community division method based on dynamic adjustment and neighbor comparison module, characterized in that: The following steps are involved: Step 1: Obtain social network data and construct the original graph G = (X, A), where A represents the adjacency matrix of the social network, representing the connection relationship between each node, X represents the feature matrix of the nodes in the social network, and the nodes represent the users in the social network; The graph enhancement module divides the original graph G = (X, A) into communities based on the sample learning positive and negative example method, and generates two new enhanced graphs G1 = (X1, A1) and G2 = (X2, A2). The specific method is: randomly delete unimportant edges in the original graph to obtain A1 and A2, and randomly delete unimportant attributes in the original graph to obtain X1 and X2; Step 2: Input the enhanced graph G1 = (X1, A1) and the enhanced graph G2 = (X2, A2) into the shared parameter two-layer encoder f θ , through the encoder f θ Perform feature extraction on the two enhanced graphs respectively to obtain the corresponding node features Z 1 and Z 2 ; Step 3: The feature Z obtained in step 2 1 and Z 2 Input adaptive adjustment module, feature Z 1 and Z 2 Dynamically adjust and reconstruct them respectively to obtain new adjacency matrices A′1 and A′2; compare A′1 with A1 and A2 with A′2 to obtain the corresponding cross entropy losses k1 and k2; the expressions are as follows: k1=[-log(A′ 1po )-log(1-A′ 1neg )] k2=[-log(A′ 2pos )-log(1-A′ 2neg )] In the above formula, T refers to the matrix transpose, σ is the Sigmoid function, and A′ 1pos Refers to A'1, A' 1neg A′ is generated by negative sampling 1pos The edge that does not exist in the view; node i represents the anchor point, and node j refers to the first-order neighbor node of node i in the view; It refers to enhancing the features of node i in graph G1. It refers to enhancing the node features of node j connected to node i in graph G1; It refers to enhancing the features of node i in graph G2. It refers to enhancing the node features of node j connected to node i in graph G2; Step 4: Use the neighbor comparison module to evaluate the contribution of each neighbor to the loss calculation and calculate the comparison loss L N When , the similarity of the adjacent nodes of the anchor point is introduced and added to the positive sample pair; the positive nodes include the same nodes of the anchor points in different views and the neighbor nodes of the anchors within the view and across different views; the negative nodes include the non-neighbors of the anchors within the view and across different views; L N The calculation formula is as follows: In the above formula, It refers to the similarity of the adjacent nodes of the anchor point; Finally, the total loss function is used to evaluate the difference between the model's predicted value and the true value; the total loss includes the contrast loss L N , reconstruction loss k1 and reconstruction loss k2; the reconstruction loss is used to minimize the difference between the reconstructed adjacency matrix and the original adjacency matrix.
2. The network community division method based on dynamic adjustment and neighbor comparison module according to claim 1 is characterized in that: The detailed method of step 1 is: Step 1.1: construct an original graph based on social network data, and divide the original graph G = (X, A) into different communities according to the community detection algorithm; Step 1.2: For the adjacency matrix A, first delete the edges between communities, then delete the edges of communities with weak community strength S, and obtain new adjacency matrices A1 and A2; the method is to use the scoring function W(e) to calculate each edge e∈ε in community c c The weight w e ; H∈{0,1} n×c is the community indicator matrix, H i,c =1[v i ∈c] represents which community the i-th node belongs to; The edge mask me is obtained by independent sampling from two Bernoulli distributions, and the edge masks me are used to determine which edges are discarded. and is a hyperparameter, which will eventually lead to the final adjacency matrices A1 and A2; Step 1.3: For the feature matrix X, remove the attributes with little influence to obtain new feature matrices X1 and X2; the method is: First, define a W a Represents the probability of removing the attribute of each node in the social network: S refers to the community strength, and abs is the absolute value to ensure it is non-negative; It is a one-dimensional normalization operation; Then, by W a Compute the attribute mask m a : in and It is a hyperparameter, ranging from 0 to 1, unequal, and then determined into two different attribute matrices X1 and X2; and Step 1.4: After the edges and attributes are changed, the original graph G = (X, A) generates two new enhanced graphs G1 = (X1, A1) and G2 = (X2, A2).
3. The network community division method based on dynamic adjustment and neighbor comparison module according to claim 2 is characterized in that: The community strength S c The calculation formula is as follows: Among them, ε is the sum of all the edges in the original graph, ε c is the sum of all edges in the community, v is the node in the original graph, and d(v) is the sum of the degrees of the nodes. The definition of strong and weak communities is to compare S c size.
4. The network community division method based on dynamic adjustment and neighbor comparison module according to claim 1 is characterized in that: The encoder f θ Based on a two-layer graph convolutional neural network GCN, sharing the same parameters; Encoder f θ The structure is as follows: Where Z (1) It represents the feature vector of the node at layer l, W (1) represents the parameters of the first layer of convolution, represents the adjacency matrix of the feature matrix, represents the degree matrix of the adjacency matrix, I represents the identity matrix, σ represents the activation function, represents the Laplacian matrix; Finally, the feature representation of the node is as follows; like Indicates that there is an edge between node i and node j; like Indicates that there is no edge between node i and node j.
Citation Information
Patent Citations
Community division method and system based on social network, and storage medium
CN113407784A
Community discovery method based on variational graph embedding
CN117056763A