Edge missing network community detection method based on dual-channel deep embedding clustering

Through the dual-channel deep embedding clustering method, combined with edge deletion and enhancement algorithm, the problem of community detection performance degradation in edge deletion networks is solved, achieving higher robustness and accuracy.

CN120296449APending Publication Date: 2025-07-11XIAN UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510334240.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing community detection algorithms have degraded or failed in edge-deleted networks, making it difficult to effectively identify community structures.

Method used

The method based on dual-channel deep embedding clustering is adopted, and edge deletion is simulated through three edge deletion strategies, and edge information is restored by combining the random walk edge enhancement algorithm. The dual-channel deep embedding clustering model is used to process core node information and higher-order neighbor information respectively, and community division is optimized through dynamic weighted fusion and self-supervised clustering.

Benefits of technology

It significantly improves the robustness and accuracy of community detection in edge-loss scenarios, can better reflect the network structure, and generate more accurate community division results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296449A_ABST
    Figure CN120296449A_ABST
Patent Text Reader

Abstract

The invention provides an edge deletion network community detection method based on dual-channel deep embedding clustering, and the method comprises the steps: firstly proposing three edge deletion strategies which are used for simulating an edge deletion condition possibly occurring in an actual network; then, an edge enhancement algorithm based on random walk is put forward to recover missing edge information under different edge deletion strategies; a clustering model based on double-channel depth embedding is introduced and comprises an information embedding core module and a self-supervised clustering core module. According to the method, the missing edge condition in a real network is simulated by providing three different edge deletion strategies, and the missing edge is effectively recovered in combination with an edge enhancement algorithm based on random walk, so that the robustness and accuracy of community detection in an edge missing scene are greatly improved. Network core node information and high-order neighbor information are respectively subjected to embedded representation, node features with higher distinction degree are obtained through a dynamic weighted fusion strategy, and more sufficient information support is provided for subsequent community division.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of community detection, and particularly relates to a community detection method for edge-missing networks based on dual-channel deep embedding clustering. Background Art

[0002] With the rapid development of the digital age, a large amount of unlabeled data has been accumulated in fields such as the Internet and social media. To better understand and utilize this data, it is particularly crucial to explore efficient analysis methods. Due to the significant local connectivity characteristics of graph data, it has become a powerful tool for extracting multi-scale local information. As an important means, community detection can reveal the potential patterns and structures in the dataset. A community is a special sub-network structure, usually referring to a set of nodes with similar characteristics. In complex networks, there may be multiple communities, and these communities have the following characteristics: the connections between nodes within a community are dense, while the connections between nodes in different communities are relatively sparse. As one of the core tasks of unsupervised learning, the goal of community detection is to divide data points into different clusters, ensuring high similarity within the clusters and relative independence between the clusters.

[0003] The research on community detection is of great significance and has been widely applied in fields such as social networks, biological networks, and citation networks. For example, in social networks, it can identify social groups; in biological networks, especially in protein-protein interaction (PPI) networks, it can reveal protein modules with similar functions; in citation networks, community detection can be used to identify research topics, explore their relationships, analyze the evolution process, and predict research trends. However, due to factors such as data collection limitations, privacy protection, and random noise, obvious edge information is often missing in actual networks. This defect in the topological structure poses a huge challenge to the community detection task.

[0004] To address the above problems, we propose a dual-channel deep embedding clustering community detection method for edge-missing networks to better perform community detection on edge-missing networks. Summary of the Invention

[0005] The purpose of the present invention is to provide a community detection method for edge-missing networks based on dual-channel deep embedding clustering, which solves the problem that traditional community detection algorithms may experience performance degradation or method failure.

[0006] The technical solution adopted by the present invention is as follows: a method for detecting edge-missing network communities based on dual-channel deep embedded clustering. First, three edge deletion strategies are proposed to simulate the possible edge-missing situations in the actual network. Subsequently, the present invention proposes an edge enhancement algorithm based on random walk to recover the missing edge information under different edge deletion strategies. Finally, a dual-channel deep embedded clustering model is introduced, including two core modules: information embedding and self-supervised clustering. In information embedding, two VAE channels are respectively used to process the core node information and high-order neighbor information of the network, and the embeddings of the two channels are dynamically weighted and fused as the input of the clustering module. The information embedding and clustering modules are jointly optimized to obtain the final community division, which specifically includes the following steps. The specific operation steps are as follows:

[0007] Step 1: Preprocess the real blog network dataset to obtain the network adjacency matrix, and use different edge deletion methods to simulate the edge information missing in different datasets;

[0008] Step 2: Use the edge enhancement method to enhance the edges of different edge-missing networks to obtain a complete network;

[0009] Step 3: Obtain the core node information and high-order neighbor information in the complete network of Step 2 and establish similarity matrices respectively;

[0010] Step 4: Use the dual-channel deep embedded clustering model to perform information embedding and clustering on the two similarity matrices in Step 3 respectively to obtain the final community division.

[0011] The characteristics of the present invention also lie in that,

[0012] Further, Step 1 specifically includes:

[0013] Step 1.1: Obtain the data in the original real blog network Polblogs dataset and establish the original network G;

[0014] Step 1.2: For the original network G in Step 1.1, adopt the random edge deletion method, randomly select the corresponding number of edges with the missing ratio according to 9 different missing rates, and remove them from the original network G to obtain the edge-missing network G′1;

[0015] For the original network G in Step 1.1, adopt the deletion method based on edge weight to obtain the edge-missing network G′2;

[0016] The deletion method based on edge weight is as follows: the sum of the degrees of the two nodes corresponding to the edge is divided by 2 as the edge weight, and then a descending order is made according to the edge weight. According to 9 different missing rates, select the corresponding number of edges with larger edge weights for deletion; the edges with larger weights usually play a more critical connection role in the network, and deleting these edges will significantly change the topological structure of the network;

[0017] For the original network G in Step 1.1, an edge deletion method based on betweenness centrality is adopted to obtain an edge-missing network G′3; specifically as follows:

[0018] First, calculate the betweenness centrality of each edge, then perform a descending order sorting according to the betweenness centrality, and select the edges with larger betweenness centrality corresponding to the missing proportion according to 9 different missing rates for deletion; the edges with high betweenness centrality are usually located at the connection points between different communities, and their deletion will weaken the connectivity between communities.

[0019] The missing rates are 0, 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, and 0.4 respectively.

[0020] Furthermore, Step 2 specifically includes:

[0021] Step 2.1: Obtain an edge-missing network G′ that deletes the number of edges with a specified missing rate;

[0022] Step 2.2: First, find all connected components {C1, C2, …, C k} in G′;

[0023] Step 2.3: Find the degree central nodes v i * in all the connected components obtained in Step 2.2;

[0024] Step 2.4: Construct a minimum spanning tree T MST for all degree central nodes;

[0025] Step 2.5: Add the edges in the minimum spanning tree T MST to the edge-missing network G′ to obtain a preliminary connected graph

[0026] Step 2.6: On the preliminary connected graph respectively, generate a path Walk(v) = {v0, v1, …, v L} with each node as the initial node v;

[0027] Step 2.7: Perform N random walks to generate a path set

[0028] Step 2.8: Count the co-occurrence frequency of any node pair (i, j) in the path set ;

[0029]

[0030] where 1 is an indicator function, which takes the value of 1 when the condition is satisfied, and 0 otherwise, is a path set;

[0031] Step 2.9: Normalize the co-occurrence frequency to obtain the co-occurrence probability of any node pair (i, j);

[0032]

[0033] where N is the number of random walks for each node, and L is the length of the random walk path.

[0034] Step 2.10: Use the normal distribution to fit the co-occurrence probability of all node pairs, and set a dynamically adjustable percentile threshold point percentile, whose value ranges from 0.995 to 0.998;

[0035] Step 2.11: Calculate the co-occurrence probability threshold threshold corresponding to this percentile according to the following formula:

[0036] threshold = Φ -1 (percentile)·std + mean;

[0037] where, Φ -1 is the quantile function of the standard normal distribution, and std and mean are respectively the standard deviation and mean of the normal distribution fitted by the co-occurrence probability co_prob(i, j) of all node pairs;

[0038] Step 2.12: Select the node pairs with co-occurrence probability greater than threshold. If there is no edge between this node pair, add it as a new edge to the connected graph and keep it if there is an edge originally, to obtain the final enhanced graph G″.

[0039] In Steps 2.10 and 2.11, the percentile of the normal distribution refers to the critical value that divides the area under the normal distribution curve according to a specific probability ratio, and the co-occurrence probability threshold refers to the corresponding normal distribution quantile calculated according to the percentile of the normal distribution;

[0040] Furthermore, Step 3 specifically includes:

[0041] Step 3.1: Calculate the core node information similarity matrix S of the edge-enhanced graph G″;

[0042] Step 3.2: Calculate the high-order neighbor information similarity matrix M of the edge-enhanced graph G″;

[0043] Further, in step 3.1, the core node information refers to the node core degree information extracted by the k-core algorithm of the network. Specifically, in a network, find the largest subgraph that satisfies that each node is connected to at least k other nodes. For each node i, its core number c(i) is defined as the k value corresponding to the largest k-core subgraph in which the node can exist; adopt a matrix construction method that combines node core degree and adjacency matrix. First, construct the adjacency matrix A of the enhanced graph G″:

[0044]

[0045] After obtaining the core number c(i) of each node, the adjacency matrix is weighted processed by the following formula:

[0046]

[0047] where V is the set of all nodes, and max(c(v)) represents the maximum value of the core numbers of all nodes in the network, which is used for normalization processing. Finally, perform a transpose flip operation on the above matrix A′ to generate the final core node similarity matrix S.

[0048] Further, in step 3.2, the definition of the high-order neighbor information similarity matrix M is:

[0049]

[0050] where T represents the transition matrix, which can be defined as:

[0051]

[0052] where, e ij represents the edge connecting node i and node j, E is the set of all edges in the graph, and d i is the degree of node i, that is, the number of edges directly connected to node i. In this paper, t = 2 is selected as the calculation range of high-order neighbor information, that is, second-order neighbor information, which not only retains the description ability of the local topological structure but also takes into account the computational efficiency.

[0053] Further, step 4 specifically includes:

[0054] Step 4.1: Construct two parallel variational autoencoder neural network models and initialize various parameters;

[0055] Step 4.2: Use the core node information similarity matrix S and the high-order neighbor information similarity matrix M obtained in step 3 as the inputs of the encoder modules in the two variational autoencoders respectively;

[0056] Step 4.3: The encoder modules of the two variational autoencoders respectively perform dimensionality reduction and feature extraction on the inputs in Step 4.2 through L-1 fully connected layers plus activation functions to obtain low-dimensional information matrices and Then, a fully connected layer is used to obtain the latent mean vector μ S , μ M and the standard deviation (or log variance) vector σ S , σ M ;

[0057] Step 4.4: Respectively perform reparameterization sampling on the mean vector μ S , μ M and the standard deviation vector σ S , σ M obtained in Step 4.3 to obtain the embedded information Z S and Z M ;

[0058] Specifically, reparameterization sampling is performed through a standard normal distribution to generate the embedded information, and the embedded information z is expressed as:

[0059]

[0060] where ⊙ represents element-wise multiplication; ∈ is a random noise vector sampled from the standard normal distribution, 0 represents a d-dimensional zero vector, and I is a d×d-dimensional identity matrix;

[0061] Step 4.5: Introduce the KL divergence loss function to measure the similarity between the latent distribution q(z|X) output by the encoder and the prior distribution p(z);

[0062] Step 4.6: Respectively use the two pieces of embedded information in Step 4.5 as the inputs of the decoder modules of the two variational autoencoders, and perform information reconstruction on the two inputs through L fully connected layers respectively to obtain the reconstruction core node information similarity matrix and the node high-order neighbor information similarity matrix

[0063] Step 4.7: Introduce the reconstruction loss function to measure the difference between the reconstructed features generated by the decoder and the original input features, and ensure that the model captures the key information of the original input in the latent space;

[0064] Step 4.8: For the two pieces of embedded information obtained in Step 4.4, two learnable weight parameters γ1 and γ2 are introduced, and the two pieces of embedded information are dynamically weighted and fused to obtain the final embedded information:

[0065] Z = γ1·ZS + γ2·Z M ;

[0066] Step 4.9: Establish a self-supervised clustering module based on the GMM single hidden layer autoencoder. Use the final embedded information Z in Step 4.8 as the input of the clustering module encoder, and finally map it to the clustering probability distribution which is the final community division result;

[0067] Step 4.10: The decoder layer of the clustering module maps the clustering probability distribution Γ back to the reconstructed number in the input space With this design, the clustering module can directly learn and optimize the clustering centers and assignment probabilities;

[0068] Step 4.11: The loss function of the clustering module is derived from the log-likelihood function of the GMM;

[0069] Step 4.12: Embed the self-supervised clustering module into the variational autoencoder to jointly optimize the embedded representation and clustering through the shared embedding space.

[0070] Furthermore, in Step 4.3, after passing through L - 1 encoder layers, and are mapped to the latent space to generate the mean and standard deviation. The calculation formulas are as follows:

[0071] μ = h (L-1) W μ + b μ , logσ 2 = h (L-1) W σ + b σ ;

[0072] W μ 、b μ 、W σ 、b σ are the weight parameters corresponding to different information respectively;

[0073] Furthermore, in Step 4.5, assume that the prior distribution p(z) is a standard normal distribution The latent distribution output by the encoder is parameterized as a Gaussian distribution by the mean μ and variance σ 2 When embedding information, the KL loss function is used to constrain the latent distribution to be close to the standard normal distribution for both views, thus stabilizing the generated latent representation. The KL loss function is defined as:

[0074]

[0075] ​Further, in step 4.7, the reconstruction loss functions of the core node similarity information matrix and the node high-order neighbor information matrix are defined as:

[0076]

[0077] where cross_entropy_loss represents the cross-entropy loss function, which is used to calculate the similarity between the original information and the reconstructed information;

[0078] Further, in step 4.10, the definition of the encoder layer Γ function of the clustering module is:

[0079] Γ = F(Z) = softmax(ZW enc + b enc )

[0080] where, and are the weight and bias of the encoder respectively, d is the dimension of the encoder embedded information, K is the number of clusters of the network, and Γ is the soft clustering matrix.

[0081] Further, in step 4.110, the decoder layer of the clustering module maps the clustering probability distribution Γ back to the reconstructed data in the input space

[0082]

[0083] where, and are the weight and bias of the decoder respectively.

[0084] Further, in step 4.11, the loss function of the clustering module is derived from the log-likelihood function of GMM and is obtained in the following form through equivalent derivation:

[0085]

[0086] where, is the reconstruction error term, which ensures that the input data can be reconstructed by the decoder; γ ik (1 - γ ik )‖μ k ‖ 2 is the sparsity regularization term, which encourages the assignment probability of data points to a certain cluster to be close to 1; is the cluster center separation term, which avoids the centers of different clusters being too close; is the balance term, which is introduced by the Dirichlet prior and is used to prevent the cluster assignment from being too biased towards a few clusters.

[0087] Furthermore, in step 4.12, two hyperparameters are introduced to balance the weights between the information embedding and the clustering module, so as to achieve the joint optimization between the two modules. The calculation formula is as follows:

[0088]

[0089] Among them, λ1 and λ2 are balance hyperparameters. The output of the encoder is used as both the input of the clustering module and for the reconstruction of the decoder. The clustering module optimizes its loss function At the same time, the parameters of the variational autoencoder are updated through gradient backpropagation, thereby enhancing the support of the embedding space for the clustering task.

[0090] The beneficial effects of the present invention are as follows:

[0091] Enhanced robustness: By proposing three different edge deletion strategies to simulate the missing edge situation in the real network and combining with the edge enhancement algorithm based on random walk to effectively recover the missing edges, the robustness and accuracy of community detection in the edge missing scenario are greatly improved.

[0092] Sufficient information fusion: The present invention uses the "dual-channel" variational autoencoder model to respectively embed and represent the network core node information and the high-order neighbor information, and obtains more discriminative node features through the dynamic weighted fusion strategy, providing more sufficient information support for subsequent community division.

[0093] Efficient self-supervised clustering: The low-dimensional features of the two channels are clustered in a self-supervised manner. By jointly optimizing the KL divergence loss and the reconstruction loss, it is ensured that the representation in the latent space can better reflect the community structure. Description of the Drawings

[0094] Figure 1 is the flowchart of the method for detecting communities in an edge missing network based on dual-channel deep embedding clustering of the present invention;

[0095] Figure 2(a) is a schematic structural diagram of randomly deleting edges;

[0096] Figure 2(b) is a schematic structural diagram of deleting edges based on edge weights. The data on each edge in the figure is the weight of the edge;

[0097] Figure 2(c) is a schematic structural diagram of deleting edges based on edge betweenness. The data on each edge in the figure is the edge betweenness value of the edge.

[0098] Figure 3 is the core structure information diagram in the k-core algorithm under different k values of the present invention.

[0099] Figure 4(a) is an example diagram of the encoder module of the variational autoencoder neural network in the deep learning module of the present invention;

[0100] Figure 4(b) is an example diagram of the self-supervised clustering module of the present invention;

[0101] Figure 5(a) - Figure 5(b) It is a comparison of community partitioning performance on the Polblogs network based on the method of randomly deleting edges. The abscissa of Figure 5(a) is the proportion of missing edges in the real network Polblogs (range: 0% - 40%), and the ordinate is the NMI value;

[0102] The abscissa of Figure 5(b) is the missing proportion of the real network Polblogs (range: 0% - 40%), and the ordinate is the ACC value;

[0103] Figure 6(a) - Figure 6(b) It is a comparison of community partitioning performance on the Polblogs network based on the edge deletion method of edge weights. The abscissa of Figure 6(a) is the proportion of missing edges in the real network Polblogs (range: 0% - 40%), and the ordinate is the NMI value; the abscissa of Figure 6(b) is the missing proportion of the real network Polblogs (range: 0% - 40%), and the ordinate is the ACC value.

[0104] Figure 7(a) - Figure 7(b) It is a comparison of community partitioning performance on the Polblogs network based on the edge deletion method of edge betweenness; the abscissa of Figure 7(a) is the proportion of missing edges in the real network Polblogs (range: 0% - 40%), and the ordinate is the NMI value; the abscissa of Figure 7(b) is the missing proportion of the real network Polblogs (range: 0% - 40%), and the ordinate is the ACC value. Detailed implementation manners

[0105] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0106] Embodiment 1

[0107] The method for detecting communities in an edge-missing network based on dual-channel deep embedding clustering of the present invention specifically comprises the following steps:

[0108] Step 1: Preprocess the real blog network dataset to obtain a network adjacency matrix, and use different edge deletion methods to simulate the missing edge information of different datasets;

[0109] Step 2: Use the edge enhancement method to enhance the edges of different edge-missing networks to obtain a complete network;

[0110] Step 3: Obtain the core node information and high-order neighbor information in the complete network obtained in Step 2 and establish similarity matrices respectively;

[0111] Step 4: Use the dual-channel deep embedding clustering model to perform information embedding and clustering on the two similarity matrices in Step 3 respectively to obtain the final community partitioning.

[0112] Example 2

[0113] The present invention proposes a method for detecting edge - missing network communities based on dual - channel deep embedding clustering. As Figure 1 shown, it mainly includes the following steps:

[0114] First, three edge - deletion strategies are proposed to simulate the possible edge - missing situations in the actual network. Subsequently, the present invention proposes a random - walk - based edge - enhancement algorithm to recover the missing edge information under different edge - deletion strategies. Finally, the present invention introduces a dual - channel deep embedding clustering model, including two core modules: information embedding and self - supervised clustering. In information embedding, two VAE channels are respectively used to process the core node information and high - order neighbor information of the network, and the embeddings of the two channels are dynamically weighted and fused as the input of the clustering module. The information embedding and clustering modules are jointly optimized to obtain the final community partition. The specific steps are as follows:

[0115] Step 1, simulation of the edge - deletion module: In order to intervene in the topological structure of the network from different perspectives, the possible edge - loss mechanisms in the real world are simulated from three different perspectives. The random uniform deletion method assumes that the edges in the network are evenly distributed, that is, all edges have the same probability of being deleted. The deletion method based on edge weights defines an importance index for edges, assigns different weight values to edges, and selects the edges with larger weights for deletion. Edge betweenness is used to evaluate the number of shortest paths passing through the edge. Edges with high betweenness are usually located at the connections between different communities, and their deletion will weaken the connectivity between communities.

[0116] Step 2, edge - enhancement module: Since the edge - missing situation will affect the performance of community detection, the present invention combines the minimum spanning tree (MST) of the graph, random walk, node similarity matrix, and Gaussian mixture model, comprehensively considering the local and global structural characteristics of the network, and gradually constructs an enhanced graph by enhancing the connectivity of the network, capturing the potential correlation information between nodes, and predicting the missing edges, thereby generating an enhanced graph and improving the performance of community detection.

[0117] Step 3, information feature matrix extraction module: In order to reduce the influence of topological structure information on community partition, the core node information is used to establish a similarity matrix, focusing on the node structure in the target set network rather than the influence of edges on the network. The correlation between nodes not only depends on the directly connected edges but also is affected by the multi - order neighbor structure. High - order neighbor information can more comprehensively describe the potential correlation between nodes, and a high - order neighbor information matrix is used to quantify the high - order topological correlation between nodes.

[0118] Step 4, Dual-channel Deep Embedding Clustering Module: Two parallel VAE models are adopted, which achieve channel-separated processing of core node information and high-order neighbor information in the structural design. By dynamically weighted fusion of embedding features at different levels, the design of the dual-channel architecture fully balances the local feature and global feature expression capabilities of the network, and can comprehensively capture the multi-level structural characteristics of the network. At the same time, a self-supervised clustering mechanism is introduced in the model optimization, and a joint optimization objective of information embedding and clustering module is constructed, thereby enhancing the representation ability of the embedding space for community structure.

[0119] Example 3

[0120] Based on Example 2, as shown in Figure 2, Step 1 specifically includes the following steps:

[0121] Step 1.1: As shown in Figure 2(a), it is random edge deletion. Assume that the edges in the network are uniformly distributed, that is, all edges have the same probability of being deleted, without considering the importance or other characteristics of the edges in the network topology. Specifically, by setting a missing rate (such as 20%), several edges are randomly selected from the network and deleted. The advantage of this method is simplicity and unbiasedness, and it is often used as a benchmark for evaluating algorithm performance.

[0122] Step 1.2: As shown in Figure 2(b), it is edge deletion based on edge weight. The deletion method based on edge weight defines an importance index for edges, assigns different weight values to edges, and selects the edges with larger weights for deletion. Specifically, the average value of the degrees of the two end nodes of the edge is used as the weight of the edge. This method assumes that the edges between high-degree nodes usually play a more critical connection role in the network, so deleting these edges will significantly change the network topology;

[0123] Step 1.3: As shown in Figure 2(c), it is edge deletion based on edge betweenness. Edge betweenness is an important index to measure the importance of an edge in the network, and is used to evaluate the number of shortest paths passing through this edge. Specifically, the betweenness of edge (i,j) is defined as:

[0124]

[0125] where, σ st represents the number of shortest paths between node s and node t, and σ st (i,j) represents the number of these shortest paths passing through edge (i,j). Edges with high betweenness are usually located at the connection points between different communities, and deleting them will weaken the connectivity between communities.

[0126] Step 2 specifically includes the following steps:

[0127] Step 2.1: First, obtain an edge-missing network;

[0128] Step 2.2: Find all connected components in the edge missing network and the degree central nodes of each connected component;

[0129] Step 2.3: Construct a minimum spanning tree from all degree central nodes;

[0130] Step 2.4: Add the edges of the minimum spanning tree to the missing edge network to obtain a preliminary connected graph G';

[0131] Step 2.5: On the connected graph G′, starting from each node as the initial node, perform N random walks with a distance of L;

[0132] Step 2.6: Based on the set of random walk paths obtained in Step 2.5, calculate the co-occurrence frequency of each node pair, and calculate the co-occurrence probability of each node pair through the co-occurrence frequency;

[0133] Step 2.7: Use a normal distribution to fit the co-occurrence probabilities of all node pairs;

[0134] Step 2.8: Based on Step 2.7, select 0.995 as the percentile threshold, calculate its corresponding probability score, select all node pairs with co-occurrence probabilities greater than the score threshold, and add them to the connected graph G′ to obtain the final edge enhanced graph.

[0135] Example 4

[0136] On the basis of Example 3, Step 3 specifically includes the following steps.

[0137] Step 3.1: Calculate the core node information similarity matrix. The core node information refers to the node core degree information extracted by the k-core algorithm of the network. Specifically, in a network, find the largest subgraph that satisfies that each node is connected to at least k other nodes; as Figure 3 shown, it shows the core structure information diagram in the k-core algorithm under different k values; for each node i, its core number c(i) is defined as the k value corresponding to the largest k-core subgraph in which the node can exist; adopt a matrix construction method that combines the node core degree and the adjacency matrix. First, construct the adjacency matrix A of the enhanced graph G″:

[0138]

[0139] After obtaining the core number c(i) of each node, the adjacency matrix is weighted through the following formula:

[0140]

[0141] Among them, \(V\) is the set of all nodes, and \(\max(c(v))\) represents the maximum value of the core numbers of all nodes in the network, which is used for normalization. Finally, a transpose flipping operation is performed on the above matrix \(A'\) to generate the final core node information similarity matrix \(S\).

[0142] Step 3.2: Calculate the high-order neighbor information similarity matrix \(M\), which is defined as:

[0143]

[0144] Among them, \(T\) represents the transition matrix, which can be defined as:

[0145]

[0146] Among them, \(e\) ij represents the edge connecting node \(i\) and node \(j\), \(E\) is the set of all edges in the graph, and \(d\) i is the degree of node \(i\), that is, the number of edges directly connected to node \(i\). In this paper, \(t = 2\) is selected as the calculation range of high-order neighbor information, that is, second-order neighbor information, which not only retains the description ability of the local topological structure but also takes into account the computational efficiency.

[0147] Embodiment 5

[0148] Based on Embodiment 4, the information embedding module and the self-supervised clustering module in Step 4 include the following steps.

[0149] As shown in Fig. 4(a), it describes the information embedding module based on the dual VAE channels. The specific steps are as follows:

[0150] Step 4.1: Construct two parallel variational autoencoder neural network models and initialize various parameters;

[0151] Step 4.2: Take the core node information similarity matrix \(S\) and the high-order neighbor information similarity matrix \(M\) obtained in Step 3 as the inputs of the encoder modules in the two variational autoencoders respectively;

[0152] Step 4.3: The encoder modules of the two variational autoencoders respectively perform dimensionality reduction and feature extraction on the inputs in Step 4.2 through \(L - 1\) fully connected layers plus activation functions to obtain the low-dimensional information matrices and After that, a fully connected layer is used to obtain the latent mean vector \(\mu\) S , \(\mu\) M and the standard deviation (or log variance) vector \(\sigma\) S , \(\sigma\) M ;

[0153] Step 4.4: To ensure the differentiability of the gradient in the backpropagation process, the reparameterization technique is adopted to combine the standard deviation vector σ and the mean vector μ, and a standard normal distribution is used for reparameterized sampling to generate the latent representation z:

[0154]

[0155] where ⊙ represents element-wise multiplication; ∈ is a random noise vector sampled from the standard normal distribution, 0 represents the d-dimensional zero vector, and I is the d×d identity matrix; the mean vector μ S , μ M and the standard deviation vector σ S , σ M obtained in Step 4.3 are respectively subjected to reparameterized sampling to obtain the embedding information Z S and Z M ;

[0156] Step 4.5: Introduce the KL divergence loss function to measure the similarity between the latent distribution q(z|X) output by the encoder and the prior distribution p(z);

[0157] Step 4.6: Respectively use the two embedding information in Step 4.5 as the inputs of the decoder modules of the two variational autoencoders, and perform information reconstruction on the two inputs through L fully connected layers respectively to obtain the similarity matrix of the reconstructed core node information and the similarity matrix of the high-order neighbor information of the nodes

[0158] Step 4.7: Introduce the reconstruction loss function to measure the difference between the reconstructed features generated by the decoder and the original input features, and ensure that the model captures the key information of the original input in the latent space;

[0159] Step 4.8: For the two embedding information obtained in Step 4.4, two learnable weight parameters γ1 and γ2 are introduced, and the two embedding information are dynamically weighted and fused to obtain the final embedding information:

[0160] Z = γ1·Z S + γ2·Z M ;

[0161] As shown in Figure 4(b), the self-supervised clustering module is described, and the specific steps are as follows:

[0162] Step 4.9: Establish a self-supervised clustering module based on the single-hidden-layer autoencoder of GMM, use the final embedding information Z in Step 4.9 as the input of the encoder of the clustering module, and finally map it to the clustering probability distribution which is the final community division result;

[0163] Step 4.10: The decoder layer of the clustering module maps the clustering probability distribution Γ back to the reconstructed number in the input space. With this design, the clustering module can directly learn and optimize the clustering centers and assignment probabilities.

[0164] Step 4.11: The loss function of the clustering module is derived from the log-likelihood function of the GMM.

[0165] Step 4.12: Embed the self-supervised clustering module into the variational autoencoder to jointly optimize the embedding representation and clustering through a shared embedding space.

[0166] Furthermore, in Step 4.3, after L-1 encoder layers, and are mapped to the latent space to generate the mean and standard deviation, and the calculation formulas are as follows:

[0167] μ = h (L-1) W μ + b μ , logσ 2 = h (L-1) W σ + b σ ;

[0168] W μ , b μ , W σ , b σ are the weight parameters corresponding to different information calculations respectively.

[0169] Furthermore, in Step 4.5, assume that the prior distribution p(z) is a standard normal distribution The latent distribution output by the encoder is parameterized as a Gaussian distribution by the mean μ and variance σ 2 When embedding information, for the two views, the KL loss function is used to constrain the latent distribution to be close to the standard normal distribution, thereby stabilizing the generated latent representation. The KL loss function is defined as:

[0170]

[0171]

[0171] Furthermore, in Step 4.8, the reconstruction loss functions of the core node similarity information matrix and the node high-order neighbor information matrix are defined as:

[0172]

[0173] Among them, cross_entropy_loss represents the cross-entropy loss function, which is used to calculate the similarity between the original information and the reconstructed information.

[0174] Further, in step 4.10, the definition of the encoder layer Γ function of the clustering module is as follows:

[0175] Γ = F(Z) = softmax(ZW enc + b enc )

[0176] Wherein, and are the weight and bias of the encoder respectively, d is the dimension of the encoder-embedded information, K is the number of clusters of the network, and Γ is the soft clustering matrix.

[0177] Further, in step 4.11, the decoder layer of the clustering module maps the clustering probability distribution Γ back to the reconstructed data in the input space

[0178]

[0179] Wherein, and are the weight and bias of the decoder respectively.

[0180] Further, in step 4.12, the loss function of the clustering module is derived from the log-likelihood function of GMM and is obtained in the following form through equivalent derivation:

[0181]

[0182] Wherein, is the reconstruction error term to ensure that the input data can be reconstructed by the decoder; γ ik (1 - γ ik )‖μ k ‖ 2 is the sparsity regularization term to encourage the assignment probability of data points to a certain cluster to be close to 1; is the cluster center separation term to avoid the centers of different clusters being too close; is the balance term introduced by the Dirichlet prior to prevent the cluster assignment from being too biased towards a few clusters.

[0183] Further, in step 4.13, two hyperparameters are introduced to balance the weights between the information embedding and the clustering module to achieve the joint optimization between the two modules. The calculation formula is as follows:

[0184]

[0185] Where λ1 and λ2 are the balance hyperparameters, and the output of the encoder is used as both the input of the clustering module and for the reconstruction of the decoder. The clustering module optimizes its loss function Meanwhile, update the parameters of the variational autoencoder through gradient backpropagation to enhance the support of the embedding space for the clustering task.

[0186] In reality, when the network topology is incomplete, the performance of community detection may decrease significantly or even fail. Existing missing edge prediction methods usually rely on strong prior assumptions, only consider local topological structure information, or only focus on global structure information; moreover, the cooperative relationship between core nodes and neighbor structures cannot be fully captured in traditional community detection; in addition, the embedding of the original input information and the clustering task are often not jointly considered. The present invention focuses on solving existing problems and proposes a method for community detection of edge-missing networks based on dual-channel deep embedding clustering. After multiple rounds of experimental verification, this method performs excellently in common evaluation indicators for community detection such as accuracy (ACC) and normalized mutual information (NMI), and has more advantages compared with other methods, and can generate complete and relatively accurate community division results.

[0187] Example 6

[0188] Use the present invention and two current advanced community detection methods to select the real network dataset Polblogs dataset to conduct comparative experiments under three different edge deletion benchmarks respectively. The Edmot method conducts community detection by constructing hypergraph segmentation and using a motif-based edge enhancement method. The CSEA first extracts core node information using non-k-truss, then uses VAE to extract and embed information features, and finally uses k-means for unsupervised clustering. The normalized mutual information (NMI) and accuracy (ACC) are the criteria for evaluating the performance of community detection methods, and their values range from 0 to 1. The closer the value is to 1, the better the performance of the community detection method.

[0189] Figure 5(a) - Figure 5(b) The experimental results of randomly deleting different proportions of edges in the real network Polblogs are shown respectively. In Figure (a), the abscissa represents the proportion of missing edges, and the ordinate is the value of the normalized mutual information (NMI); in Figure (b), the abscissa also represents the proportion of missing edges, and the ordinate is the value of the accuracy (ACC). The experimental results show that on the Polblogs network, the community detection performance of the present invention is better than that of the other two advanced community detection methods when dealing with real missing edge networks.

[0190] Figure 6(a) - Figure 6(b)The experimental results of deleting different proportions of edges based on edge weights in the real network Polblogs are respectively shown. In Fig. (a), the abscissa represents the proportion of missing edges, and the ordinate is the value of normalized mutual information (NMI); in Fig. (b), the abscissa also represents the proportion of missing edges, and the ordinate is the value of accuracy (ACC). The experimental results show that in the Polblogs network, the community detection performance of the present invention in dealing with real missing-edge networks is better than that of the other two advanced community detection methods.

[0191] Figure 7(a) - Figure 7(b) The experimental results of deleting different proportions of edges based on edge betweenness in the real network Polblogs are respectively shown. In Fig. (a), the abscissa represents the proportion of missing edges, and the ordinate is the value of normalized mutual information (NMI); in Fig. (b), the abscissa also represents the proportion of missing edges, and the ordinate is the value of accuracy (ACC). The experimental results show that in the Polblogs network, the community detection performance of the present invention in dealing with real missing-edge networks is better than that of the other two advanced community detection methods.

Claims

1. Edge-missing network community detection method based on dual-channel deep embedded clustering, characterized in that The specific steps are as follows: Step 1: Preprocess the real blog network dataset to obtain a network adjacency matrix, and use different edge deletion methods to simulate the missing edge information of different datasets; Step 2: Use the edge enhancement method to enhance the edges of different edge-missing networks to obtain a complete network; Step 3: Obtain the core node information and high-order neighbor information in the complete network of Step 2 and establish similarity matrices respectively; Step 4: Use the dual-channel deep embedding clustering model to perform information embedding and clustering on the two similarity matrices in Step 3 respectively to obtain the final community division.

2. The method for edge-missing network community detection based on dual-channel deep embedding clustering according to claim 1, wherein The specific content of Step 1 is as follows: Step 1.1: Obtain the data in the original real blog network Polblogs dataset and establish the original network G; Step 1.2: For the original network G in Step 1.1, adopt the random edge deletion method, randomly select the corresponding number of edges according to 9 different missing rates, and remove them from the original network G to obtain the edge-missing network G′1; For the original network G in Step 1.1, adopt the edge weight-based deletion method to obtain the edge-missing network G′2; The edge weight-based deletion method is as follows: The sum of the degrees of the two nodes corresponding to the edge is divided by 2 as the edge weight, and then a descending order is made according to the edge weight. According to 9 different missing rates, select the edges with larger edge weights corresponding to the missing proportion of the number of edges for deletion; For the original network G in Step 1.1, adopt the edge betweenness-based deletion method to obtain the edge-missing network G′3; specifically as follows: First, calculate the edge betweenness of each edge, then make a descending order according to the edge betweenness, and select the edges with larger edge betweenness corresponding to the missing proportion of the number of edges according to 9 different missing rates for deletion; The missing rates are 0, 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4 respectively.

3. The method for detecting edge - missing network communities based on dual - channel deep embedding clustering according to claim 2, wherein, The specific content of Step 2 includes: Step 2.1: Obtain an edge-missing network with the number of edges deleted at a specified missing rate; Step 2.2: First, find all connected components {C1, C2, …, C k} in the edge-missing network; Step 2.3: Find the degree central nodes in all connected components obtained in Step 2.2 Step 2.4: Construct a minimum spanning tree T for all degree central nodes MST ; Step 2.5: Add the edges in the minimum spanning tree T MST to the edge-missing network to obtain a preliminary connected graph Step 2.6: On the preliminary connected graph respectively, with each node as the initial node v, generate a path of length L, Walk(v) = {v0, v1, …, v L}; Step 2.7: Perform N random walks to generate a path set Step 2.8: Count the co-occurrence frequency of any node pair (i, j) in the path set : where 1 is an indicator function that takes the value 1 when the condition holds and 0 otherwise, is a set of paths; Step 2.9: Normalize the co-occurrence frequency to obtain the co-occurrence probability of any node pair (i,j): where N is the number of random walks of each node and L is the length of the random walk path; Step 2.10: Use the normal distribution to fit the co-occurrence probability of all node pairs, and set a dynamically adjustable percentile threshold percentile, where the value of percentile is 0.995 - 0.998; Step 2.11: Calculate the co-occurrence probability threshold threshold corresponding to this percentile threshold according to the following formula: threshold = Φ -1 (percentile)·std + mean; where Φ -1 is the quantile function of the standard normal distribution, and std and mean are the standard deviation and mean of the normal distribution fitted by the co-occurrence probability co_prob(i, j) of all node pairs, respectively; In Steps 2.10 and 2.11, the percentile of the normal distribution refers to the critical value that divides the area under the normal distribution curve according to the probability ratio, and the co-occurrence probability threshold refers to the corresponding normal distribution quantile calculated according to the percentile of the normal distribution; Step 2.12: Select node pairs with co-occurrence probability greater than threshold. If there is no edge between the node pair, add it as a new edge to the connected graph. If there is already an edge, keep it, and obtain the final enhanced graph G″. If there is already an edge, keep it, and obtain the final enhanced graph G″.

4. The method for edge-missing network community detection based on dual-channel deep embedding clustering according to claim 3, wherein The specific content of Step 3 includes: Step 3.1: Calculate the similarity matrix S of the core node information of the edge-enhanced graph G″; The core node information refers to the node core degree information extracted by the k-core algorithm of the network. Specifically, in a network, find the largest k-core subgraph that satisfies that each node is connected to at least k other nodes; for each node i, its core number c(i) is defined as the k value corresponding to the node that can exist in the largest k-core subgraph; adopt a matrix construction method that combines node core degree and adjacency matrix. First, construct the adjacency matrix A of the enhanced graph G″: Among them, i and j respectively refer to the nodes, and E represents the edge set; After obtaining the core number c(i) of each node, perform weighted processing on the adjacency matrix through the following formula: Among them, V is the set of all nodes, and max(c(v)) represents the maximum value of the core numbers of all nodes in the network, which is used for normalization processing; finally, perform a transpose flip operation on the matrix A′ to generate the final core node information similarity matrix; Step 3.2: Calculate the high-order neighbor information similarity matrix M of the edge enhanced graph G″; The definition of the high-order neighbor information similarity matrix M is: Among them, T represents the transition matrix, and the definition is: Among them, e ij represents the edge connecting node i and node j, and E is the set of all edges in the graph. d i is the degree of node i, that is, the number of edges directly connected to node i; select t = 2 as the calculation range of the high-order neighbor information, that is, the second-order neighbor information.

5. The method for edge - missing network community detection based on dual - channel deep embedding clustering according to claim 4, wherein, Step 4 is specifically as follows: Step 4.1: Construct two parallel variational autoencoder neural network models and initialize various parameters; Step 4.2: Use the core node information similarity matrix S and the high-order neighbor information similarity matrix M obtained in Step 3 as the inputs of the encoder modules in the two variational autoencoders respectively; Step 4.3: The encoder modules of the two variational autoencoders respectively perform dimensionality reduction and extraction of features on the inputs in Step 4.2 through L-1 fully connected layers plus activation functions, obtaining low-dimensional information matrices and Then, a fully connected layer is used to obtain the latent mean vector μ S , μ M and the standard deviation vector σ S , σ M ; Step 4.4: Respectively perform reparameterization sampling on the mean vectors μ S , μ M and the standard deviation vectors σ S , σ M to obtain the embedding information Z S and Z M ; Specifically, through a standard normal distribution reparameterized sampling is performed to generate embedding information, and the embedding information z is expressed as: Among them, ⊙ represents element-wise multiplication; ∈ is a random noise vector sampled from the standard normal distribution, 0 represents a d-dimensional zero vector, and I is a d×d-dimensional identity matrix; Step 4.5: Introduce the KL divergence loss function which is used to measure the similarity between the latent distribution q(z|X) output by the encoder module and the prior distribution p(z); Step 4.6: Use the two embedding information Z S and Z M as the inputs of the decoder modules of two variational autoencoders respectively, and reconstruct the information of the two inputs through L fully-connected layers respectively to obtain the similarity matrix of the reconstructed core node information and the similarity matrix of the high-order neighbor information of the nodes Step 4.7: Introduce the reconstruction loss function It is used to measure the difference between the reconstructed features generated by the decoder and the original input features, ensuring that the model captures the key information of the original input in the latent space; The reconstruction loss functions of the core node similarity information matrix and the node high-order neighbor information matrix are defined as: Among them, cross_entropy_loss represents the cross-entropy loss function, which is used to calculate the similarity between the original information and the reconstructed information; Step 4.8: For the two embedding information obtained in Step 4.4, introduce two learnable weight parameters γ1 and γ2, and dynamically weight and fuse the two embedding information to obtain the final embedding information: Z = γ1·Z S + γ2·Z M ; Step 4.9: Establish a self-supervised clustering module based on the GMM single-hidden-layer autoencoder. Use the final embedded information Z in Step 4.8 as the input to the encoder layer of the clustering module, and finally map it to a clustering probability distribution which is the final community partition result; Step 4.10: The decoder layer of the clustering module maps the clustering probability distribution Γ back to the reconstructed number in the input space Among them, and are the weights and biases of the decoder of the clustering module, respectively; Step 4.11: Loss function of the clustering module Derived from the log-likelihood function of GMM; the following form is obtained through equivalent derivation: Among them, is the reconstruction error term, ensuring that the input data can be reconstructed by the decoder; γ ik (1 - γ ik )‖μ k ‖ 2 is the sparsity regularization term; is the clustering center separation term; is the balance term, introduced by the Dirichlet prior, which is used to prevent the cluster assignment from being too biased towards a few clusters; Step 4.12: Embed the self-supervised clustering module into the variational autoencoder, and realize the joint optimization of the embedding representation and clustering through the shared embedding space.

6. The method for edge - missing network community detection based on dual - channel deep embedding clustering according to claim 7, wherein, Step 4.3 is specifically as follows: After passing through L-1 encoder layers, and are mapped to the latent space to generate the mean and standard deviation, and the calculation formulas are as follows: μ = h (L-1) W μ + b μ , logσ 2 = h (L-1) W σ + b σ ; W μ and b μ 、W σ and b σ are the weight parameters corresponding to different information respectively.

7. The method for edge-missing network community detection based on dual-channel deep embedding clustering according to claim 6, wherein Step 4.5 is specifically as follows: Assume that the prior distribution p(z) is a standard normal distribution The latent distribution output by the encoder is parameterized by the mean μ and the variance σ 2 as a Gaussian distribution When embedding information, the KL loss function is used for the two views respectively to constrain the latent distribution to be close to the standard normal distribution, so as to stabilize the generated latent representation; the KL loss function is defined as:

8. The method for edge - missing network community detection based on dual - channel deep embedding clustering according to claim 7, wherein, Step 4.10 is specifically as follows: The definition of the probability distribution function Γ of the encoder layer of the clustering module is: Γ = F(Z) = softmax(ZW enc + b enc ) wherein, and are the weights and biases of the encoder respectively, d is the dimension of the encoder's embedded information, K is the number of clusters of the network, and Γ is the soft clustering matrix.

9. The method for edge - missing network community detection based on dual - channel deep embedding clustering according to claim 8, wherein, In Step 4.12, introduce two hyperparameters to balance the weights between the information embedding and the clustering module to achieve the joint optimization between the two modules. The calculation formula is as follows: Among them, λ1 and λ2 are balanced hyperparameters. The output of the encoder serves as both the input to the clustering module and for reconstruction by the decoder; the clustering module optimizes its loss function Meanwhile, the parameters of the variational autoencoder are updated through backpropagation of gradients, thereby enhancing the support of the embedding space for the clustering task.

Citation Information

Cited By

  • Dynamic quantile filtering method for commodity co-occurrence network

    CN122089367A