Community discovery method and system based on text deep clustering and storage medium
By combining the attention mechanism with the bidirectional long and short-term memory network and the convolutional neural network, the text depth features are extracted and self-supervised cluster loss function is constructed, the problem of unsatisfactory community division effect in the text network is solved, and more accurate and stable community discovery results are achieved.
Patent Information
- Application Number
- CN202510029584.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art has poor community division effect in text networks, insufficient feature representation, difficult to extract correlation and contextual semantic information between words in text, and lack effective supervision signals in the community discovery process.
The bidirectional long and short-term memory network model combined with attention mechanism is used to extract the global context-dependent information of the text and the local information between continuous words, realize the deep feature fusion of the text, and build a self-supervised cluster loss function based on deep feature fusion to guide feature learning.
The accuracy and stability of community discovery results were improved, and the feature representation and community division effects were optimized through deep feature fusion and self-supervised clustering loss function.
Smart Images

Figure CN120030165A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology and network science based on computer technology, and in particular to a community discovery method and system based on text deep clustering. Background Art
[0002] In real life, many complex systems can be modeled as a complex network for analysis, such as common traffic networks, protein interaction networks, and text networks. Community discovery is an important task in network analysis. It divides the community structure according to the degree of connection between network nodes and reasonably classifies the nodes in the network. This technology has a wide range of applications in social network analysis, recommendation systems, computational biology and other fields.
[0003] Text has the characteristics of sparse semantic features and high spatial dimension. These characteristics make it difficult to extract effective text features from it, making the community division effect of traditional community discovery algorithms on text networks unsatisfactory. Deep clustering methods combine deep learning and clustering, using deep neural networks to learn feature representations that are conducive to clustering, and apply them to community discovery to improve the community division effect. However, existing community discovery methods combined with deep clustering have problems such as the quality of community division is only affected by feature representation, and it is difficult to extract the relevance and contextual semantic information between words in the text.
[0004] In addition, the existing deep clustering models lack effective supervision signal guidance in the process of community discovery through clustering, resulting in insufficient representation of text features by the model and low stability of community division results. For example, the public document with the publication number CN109299749A, the publication date February 1, 2019, and the patent name "A Community Discovery Algorithm Based on Improved Density Peak Clustering" first uses the equivalent resistance path length as a distance metric to measure the distance between nodes in a complex network; secondly, based on the generated node distance, the density peak clustering (DPC) algorithm is run to generate a decision graph; then, on the decision graph of the density peak algorithm, the cluster center is automatically selected through the DBSCAN algorithm, rather than manually selected by observing the decision graph.
[0005] The main flaws of the algorithm currently discovered by the text community are:
[0006] 1. When representing text features, only the global context dependency information of the text or the local information between consecutive words is considered, which leads to insufficient feature representation and affects the community discovery effect;
[0007] 2. Due to the lack of supervision signals in the community discovery process, it is impossible to optimize the model parameters based on the community discovery results, resulting in low model performance. Summary of the invention
[0008] The technical problem to be solved by the present invention is to realize a community discovery method based on text deep clustering, which can realize the computer's efficient and accurate recognition ability of text data.
[0009] It includes a bidirectional long short-term memory network model combined with an attention mechanism to extract the global contextual dependency information of the text, a convolutional neural network model to extract local information between consecutive words in the text, deep feature fusion of the text, and a self-supervised clustering loss function based on deep feature fusion, so that the community discovery task can guide feature learning, obtain optimized feature vector representation, and improve the community discovery results.
[0010] In order to achieve the above object, the technical solution adopted by the present invention is: a community discovery method based on text deep clustering, comprising the following steps:
[0011] Step 1: Generate word vectors through training using the training model, and convert the preprocessed text into a word matrix;
[0012] Step 2: extract the global context information of the text and the local information between consecutive words, concatenate them to obtain the text feature vector, calculate the cosine similarity of the text feature vector, and construct the text network;
[0013] Step 3: Perform community discovery based on clustering to obtain the initial community center node and community discovery results;
[0014] Step 4: Construct a deep self-supervised clustering loss function, update the text network and community center by minimizing the loss function, and obtain the updated text feature vector and community discovery results.
[0015] In step 1, the word vector is generated by using GloVe model training;
[0016] The preprocessed text set is V = {v 1 ,v 2 ,…,v n}, each word v i Use the trained word vector e i Indicates that the text is represented by E = {e 1 ,e 2 ,…,e n}, vector e 1 To e n For word v 1 to v n The corresponding word vectors are concatenated to obtain a two-dimensional matrix:
[0017]
[0018] In the formula, e iRepresents the word vector of the i-th word in the text, X∈E n×d is the vectorized representation of the text, n is the maximum length of all texts in the corpus, and each text is padded to the maximum text length, d represents the dimension of the word vector, is the concatenation operator.
[0019] In step 2, a bidirectional long short-term memory network model and a convolutional neural network model combined with an attention mechanism are used to extract the global context information of the text and the local information between consecutive words.
[0020] The step 2 first uses a bidirectional long short-term memory network model to extract contextual semantic features of the word embedding matrix, calculates the contribution of different words to the community discovery results through the attention mechanism, obtains text vectors whose contribution is greater than a set threshold, and obtains global context dependency information; at the same time, the convolution layer and pooling layer in the convolutional neural network model are used to extract local information between consecutive words in the text; then, the output of the convolutional neural network layer is spliced and fused with the output of the attention layer to obtain a text feature vector; finally, the cosine similarity of the text feature vector is calculated to construct a text network.
[0021] The step 2 comprises:
[0022] Step 2.1: In the bidirectional long short-term memory network model, calculate each word e i The above information
[0023]
[0024] In the formula, and is the weight matrix of the bidirectional long short-term memory network, is the bias term.
[0025] Calculate each word e i The following information
[0026]
[0027] Step 2.2: Combine the previous and following information of each word to calculate the context information y of all words in the text i :
[0028]
[0029] Step 2.3: Get the word context information y i As the input of the attention layer, the normalized exponential function is used to calculate the attention weight a of the i-th word in the text i :
[0030]
[0031] in,
[0032] u t =tanh(W t y i +b t )
[0033] In the formula, u t is the hidden layer vector y i The result obtained after a fully connected layer operation, W t and b t are the weight matrix and bias term of the attention mechanism respectively; L is the length of the text, u w is a randomly initialized weight vector;
[0034] Step 2.4: According to the attention weight value calculated in step 2.3, the context vector y of each word i Perform weighted summation to obtain the text vector s of the attention mechanism layer:
[0035]
[0036] Step 2.5: The text word matrix obtained in step 1 is used as the input of the convolutional neural network. Assume that the size of the convolution kernel is w∈R h×d , where h represents the window width of the filter. When the convolution kernel is applied to a window consisting of h words, a new feature c can be generated. ij , the calculation formula is as follows:
[0037] c ij =f(wX j:j+h-1 +b)
[0038] In the formula, c ij represents the jth value on the i-th feature map; i is the number of convolution kernels, j is in the range of [1:n-h+1], indicating the number of times the convolution kernel can slide when the step size is 1; f is the ReLU activation function, and b is the bias term;
[0039] Step 2.6: When the number of convolution kernels is m, all convolution kernels are convolved with the word vector matrix and the text word matrix to obtain m corresponding feature maps c = {c 1 ,c 2 ,…,c n}; The pooling layer performs convolution on each feature map c i =[c i,1 ,c i,2 ,…,c i,n-h+1 ]Perform the maximum pooling operation and select the maximum value As the most important feature in the feature map, the other feature maps are similar, and the m eigenvalues obtained are concatenated After passing through the fully connected layer, the text vector d is obtained;
[0040] Step 2.7: Concatenate the output s of the attention mechanism layer in step 2.4 and the output d of the convolutional neural network layer in step 2.6, pass them to the fully connected layer, and obtain the text feature vector z;
[0041] Step 2.8: Calculate the cosine similarity based on the text feature vector and calculate the text feature vector z i and z j The cosine similarity sim ij :
[0042]
[0043] Where ||·|| is the L2 norm of the vector;
[0044] Step 2.9: Construct a text network based on the similarity between text feature vectors, use text as nodes in the network, set a threshold θ, and if the similarity between two texts is sim ij >θ, then add an edge between the two texts;
[0045] In step 3, K-means clustering is used to discover communities; the nodes in the network are clustered to explore the potential relationships between the nodes, and the following are obtained:
[0046] Initial community center u=[u 1 ,…,u j ,...,u k ];
[0047] Community division result C = [C 1 ,…,C j ,...,C k ].
[0048] In step 4, a clustering loss function for the community discovery task is constructed based on text features and community centers, so that the community discovery task can guide feature learning. Finally, the Adam optimization function and the back propagation algorithm are combined to minimize the loss function, and the network model parameters are updated to improve the text features and community centers to obtain the final community division results.
[0049] The step 4 comprises:
[0050] Step 4.1: Construct clustering loss function LS:
[0051]
[0052] in,
[0053]
[0054] In the formula, q i =[q i1 ,...,q ij ,...,q ik ], q ij To use the Student's t-distribution function as the kernel to measure the sample point z i To the Community Center j The similarity of is the clustering soft assignment function; p i =[p i1 ,...,p ij ,...,p ik ], p ij Assign a function to the auxiliary target as the community discovery result q ij The true label of
[0055] Step 4.2: Combine the Adam optimization function and the back propagation algorithm to minimize the clustering loss function, continuously reduce the difference between the clustering soft allocation function and the target allocation function to learn the results with higher confidence in community division, and update the network model parameters to optimize the text features and community centers. After network training, the community discovery result C i is the clustering soft assignment function q ij With the maximum probability, the text feature z i Assigned to community center j The label, that is
[0056] A community discovery system based on text deep clustering includes a processor and a storage medium, wherein the storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the community discovery method based on text deep clustering is implemented.
[0057] A storage medium is a computer-readable storage medium for storing software program codes, wherein the software program codes are used to execute the community discovery method based on text deep clustering.
[0058] The present invention only considers the global context dependency information of the text or the local information between consecutive words when representing the features of the text, which leads to the problem of insufficient feature representation. A bidirectional long short-term memory network model combined with an attention mechanism is introduced to extract the global context dependency information of the text, and a convolutional neural network model is used to extract the local information between consecutive words in the text, so as to realize the deep feature fusion of the text. In view of the problem that the community discovery process lacks supervision signals and the model parameters cannot be optimized according to the community discovery results, a self-supervised clustering loss function based on deep feature fusion is constructed, so that the community discovery task can guide feature learning, obtain an optimized feature vector representation, and improve the community division result.
[0059] The above scheme uses an attention-based bidirectional long short-term memory network model and a convolutional neural network model to extract contextual semantic features and local features between consecutive words in the text, and constructs a self-supervised clustering loss function embedded in the deep learning model to simultaneously perform feature learning and community assignment. Therefore, it can solve the problem that the existing community discovery method only considers the global and local features of the text and the community discovery process cannot optimize the model parameters according to the community division results. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The following is a brief description of the contents expressed in each figure in the specification of the present invention:
[0061] Figure 1 Flow chart of the community discovery method based on text deep clustering;
[0062] Figure 2 This is a schematic diagram of the structure of a community discovery system based on deep text clustering;
[0063] Figure 3 A flowchart of the community discovery method based on deep text clustering. DETAILED DESCRIPTION
[0064] The following is a further detailed description of the specific implementation methods of the present invention, such as the shape, structure, relative position and connection relationship between the various components involved, the function and working principle of each part, the manufacturing process and operation method, etc., through the description of the embodiments with reference to the accompanying drawings, so as to help those skilled in the art to have a more complete, accurate and in-depth understanding of the inventive concept and technical solution of the present invention.
[0065] like Figure 2 As shown in FIG. 1 , the community discovery method based on text deep clustering includes the following steps:
[0066] Step 1: Use the GloVe model to train and generate word vectors, and convert the preprocessed text into a word matrix;
[0067] Step 2: Use the bidirectional long short-term memory network model combined with the attention mechanism and the convolutional neural network model to extract the global context information of the text and the local information between consecutive words, concatenate them to obtain the text feature vector, calculate the cosine similarity of the text feature vector, and construct the text network;
[0068] Step 3: Use K-means clustering to discover communities and obtain the initial community center nodes and community division results;
[0069] Step 4: Construct a deep self-supervised clustering loss function, update the network parameters and community centers by minimizing the loss function, and obtain the updated text feature vector and community discovery results.
[0070] Each step is described in detail below:
[0071] In step 1, assume that the preprocessed text set is V = {v 1 ,v 2 ,…,v n}, each word v i Use the trained word vector e i Indicates that the text is represented by E = {e 1 ,e 2 ,…,e n}, vector e 1 To e n For word v 1 to v n The corresponding word vectors are concatenated to obtain a two-dimensional matrix:
[0072]
[0073] In the formula, e i Represents the word vector of the i-th word in the text, X∈E n×d is the vectorized representation of the text, n is the maximum length of all texts in the corpus, and each text is padded to the maximum text length, d represents the dimension of the word vector, is the concatenation operator.
[0074] In step 2, first, use the bidirectional long short-term memory network model to extract the contextual semantic features of the word embedding matrix, calculate the contribution of different words to the community discovery results through the attention mechanism, and obtain the text vector with a larger contribution, so as to capture the global context dependency information; then, use the convolution layer and pooling layer in the convolutional neural network model to extract the local information between consecutive words in the text, and concatenate and fuse the output of the convolutional neural network model with the output of the attention layer to obtain the text feature vector; finally, calculate the cosine similarity of the text feature vector and construct the text network. Specifically include:
[0075] Step 2.1: In the bidirectional long short-term memory network model, calculate each word e i The above information
[0076]
[0077] In the formula, and is the weight matrix of the bidirectional long short-term memory network, is the bias term.
[0078] Calculate each word e i The following information
[0079]
[0080] Step 2.2: Combine the previous and following information of each word to calculate the context information y of all words in the text i :
[0081]
[0082] Step 2.3: Get the word context information y i As the input of the attention layer, the normalized exponential function is used to calculate the attention weight a of the i-th word in the text i :
[0083]
[0084] in,
[0085] u t =tanh(W t y i +b t )
[0086] In the formula, u t is the hidden layer vector y i The result obtained after a fully connected layer operation, W t and b t are the weight matrix and bias term of the attention mechanism respectively; L is the length of the text, u w is a randomly initialized weight vector.
[0087] Step 2.4: According to the attention weight value calculated in step 2.3, the context vector y of each word i Perform weighted summation to obtain the text vector s of the attention mechanism layer:
[0088]
[0089] Step 2.5: The text representation X calculated in step 1 is used as the input of the convolutional neural network. Assume that the size of the convolution kernel is w∈R h×d , where h represents the window width of the filter. When the convolution kernel is applied to a window consisting of h words, a new feature c can be generated. ij , the calculation formula is as follows:
[0090] c ij =f(wX j:j+h-1 +b)
[0091] In the formula, c ij represents the jth value on the i-th feature map; i is the number of convolution kernels, j is in the range of [1:n-h+1], indicating the number of times the convolution kernel can slide when the step size is 1; f is the ReLU activation function, and b is the bias term.
[0092] Step 2.6: When the number of convolution kernels is m, all convolution kernels are convolved on the word vector matrix X to obtain m corresponding feature maps c = {c 1 ,c 2 ,…,c n}. The pooling layer performs convolution on each feature map c i =[c i,1 ,c i,2 ,…,c i,n-h+1 ]Perform the maximum pooling operation, that is, select the maximum value As the most important feature in the feature map, the other feature maps are similar, and the m eigenvalues obtained are concatenated After passing through the fully connected layer, the text vector d is obtained.
[0093] Step 2.7: Concatenate the output s of the attention mechanism layer in step 2.4 and the output d of the convolutional neural network layer in step 2.6, pass them to the fully connected layer, and obtain the text feature vector z.
[0094] Step 2.8: Calculate the cosine similarity based on the text feature vector and calculate the text feature vector z i and z j The cosine similarity sim ij :
[0095]
[0096] Where ||·|| is the L2 norm of the vector.
[0097] Step 2.9: Construct a text network based on the similarity between text feature vectors, with text as nodes in the network. Set a threshold θ, if the similarity between two texts is sim ij>θ, then add an edge between the two texts.
[0098] In step 3, K-means clustering is used to discover communities in the text network, cluster the nodes in the network to explore the potential relationships between the nodes, and obtain the initial community center u = [u 1 ,…,u j ,...,u k ] and the community division result C=[C 1 ,…,C j ,…,C k ].
[0099] In step 4, a clustering loss function for community discovery tasks is constructed based on text features and community centers, so that the community discovery task can guide feature learning. Finally, the Adam optimization function and the back propagation algorithm are combined to minimize the clustering loss function, and the network model parameters are updated to improve the text features and community centers to obtain the final community division results. Specifically, it includes:
[0100] Step 4.1: Construct clustering loss function LS:
[0101]
[0102] in,
[0103]
[0104] In the formula, q i =[q i1 ,...,q ij ,...,q ik ], q ij To use the Student's t-distribution function as the kernel to measure the sample point z i To the Community Center j The similarity of is the clustering soft assignment function; p i =[p i1 ,...,p ij ,...,p ik ], p ij Assign a function to the auxiliary target as the community discovery result q ij The real label.
[0105] Step 4.2: Combine the Adam optimization function and the back propagation algorithm to minimize the clustering loss function, continuously reduce the difference between the clustering soft allocation function and the target allocation function to learn the results with higher confidence in community division, and update the network model parameters to optimize the text features and community centers. After network training, the community discovery result C i is the clustering soft assignment function q ij With the maximum probability, the text feature zi Assigned to community center j The label, that is
[0106] The community discovery method and system based on text deep clustering include a processor and a storage medium, wherein the storage medium is a computer-readable storage medium for storing software program codes, and the software program codes are used to execute the community discovery method based on text deep clustering.
[0107] like Figure 2 As shown, a structural diagram of a community discovery system based on text deep clustering is shown. For the sake of convenience, only the part related to the embodiment of the present invention is shown. The system includes:
[0108] Text dataset preprocessing module: Use the GloVe model to train a large-scale corpus to obtain the word vectors of all words in the corpus, and then convert the preprocessed text into a word matrix consisting of stacked corresponding word vectors;
[0109] Deep feature extraction module: Use a bidirectional long short-term memory network model combined with an attention mechanism and a convolutional neural network model to extract the global context dependency information of the text and the local information between consecutive words in the text, and then concatenate them to obtain a text feature vector;
[0110] Text network construction module: Calculate the similarity between text feature vectors based on cosine similarity, build a text network based on the similarity, use text as nodes in the network, and determine whether there are edge connections between nodes based on the similarity between texts;
[0111] Self-supervised clustering module: Use K-means for community discovery, build a self-supervised clustering loss function based on deep feature fusion according to text features and community center nodes, and update network parameters and community center nodes by minimizing the loss function, so that the community division task through clustering can guide feature learning and obtain optimized feature vector representation;
[0112] Community discovery module: Based on clustering, similar feature nodes are aggregated and non-similar feature nodes are discretized. After network training, the final community division result is obtained.
[0113] Text has the characteristics of sparse semantic features and high spatial dimension. These characteristics make it difficult to extract effective text features from it, making the traditional community discovery algorithm not ideal for community division in text. Deep clustering methods combine deep learning and clustering, using deep neural networks to learn feature representations that are conducive to clustering, and apply them to community discovery to improve the community division effect. However, the existing community discovery methods combined with deep clustering have problems such as the quality of community division is only affected by feature representation, and it is difficult to extract the correlation and contextual semantic information between words in the text. In addition, the existing deep clustering models lack effective supervision signal guidance in the community discovery part through clustering, resulting in insufficient feature representation of the text and low stability of community division results.
[0114] In view of the problem that in the prior art, when representing text features, only the global context dependency information of the text or the local information between consecutive words is considered, resulting in insufficient feature representation, the present invention introduces a bidirectional long short-term memory network model combined with an attention mechanism to extract the global context dependency information of the text, and uses a convolutional neural network model to extract the local information between consecutive words in the text, thereby realizing deep feature fusion of the text; in view of the problem that the community discovery process lacks supervision signals and the model parameters cannot be optimized according to the community discovery results, a self-supervised clustering loss function based on deep feature fusion is constructed, so that the community discovery task can guide feature learning, obtain an optimized feature vector representation, and improve the community division result.
[0115] The present invention is described above by way of example in conjunction with the accompanying drawings. It is obvious that the specific implementation of the present invention is not limited to the above-mentioned method. As long as various non-substantial improvements are made using the method concept and technical solution of the present invention, or the concept and technical solution of the present invention are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.
Claims
1. A community discovery method based on text deep clustering, characterized in that: The following steps are involved: Step 1: Generate word vectors through training using the training model, and convert the preprocessed text into a word matrix; Step 2: extract the global context information of the text and the local information between consecutive words, concatenate them to obtain the text feature vector, calculate the cosine similarity of the text feature vector, and construct the text network; Step 3: Perform community discovery based on clustering to obtain the initial community center node and community discovery results; Step 4: Construct a deep self-supervised clustering loss function, update the text network and community center by minimizing the loss function, and obtain the updated text feature vector and community discovery results.
2. The community discovery method according to claim 1, characterized in that: In step 1, the word vector is generated by using GloVe model training; The preprocessed text set is V = {v1, v2, ..., v n }, each word v i Use the trained word vector e i Indicates that the text is represented by E = {e1, e2, …, e n }, vector e1 to e n For words v1 to v n The corresponding word vectors are concatenated to obtain a two-dimensional matrix: In the formula, e i Represents the word vector of the i-th word in the text, X∈E n×d is the vectorized representation of the text, n is the maximum length of all texts in the corpus, and each text is padded to the maximum text length, d represents the dimension of the word vector, is the concatenation operator.
3. The community discovery method according to claim 1, characterized in that: In step 2, a bidirectional long short-term memory network model and a convolutional neural network model combined with an attention mechanism are used to extract the global context information of the text and the local information between consecutive words.
4. The community discovery method according to claim 3, characterized in that: The step 2 first uses a bidirectional long short-term memory network model to extract contextual semantic features of the word embedding matrix, calculates the contribution of different words to the community discovery results through the attention mechanism, obtains text vectors whose contribution is greater than a set threshold, and obtains global context dependency information; at the same time, the convolution layer and pooling layer in the convolutional neural network model are used to extract local information between consecutive words in the text; then, the output of the convolutional neural network layer is concatenated and fused with the output of the attention layer to obtain a text feature vector; finally, the cosine similarity of the text feature vector is calculated to construct a text network.
5. The community discovery method according to claim 4, characterized in that: The step 2 comprises: Step 2.1: In the bidirectional long short-term memory network model, calculate each word e i The above information In the formula, and is the weight matrix of the bidirectional long short-term memory network, is the bias term; Calculate each word e i The following information Step 2.2: Combine the previous and following information of each word to calculate the context information y of all words in the text i : Step 2.3: Get the word context information y i As the input of the attention layer, the normalized exponential function is used to calculate the attention weight a of the i-th word in the text i : in, u t =tanh(W t y i +b t ) In the formula, u t is the hidden layer vector y i The result obtained after a fully connected layer operation, W t and b t are the weight matrix and bias term of the attention mechanism respectively; L is the length of the text, u w is a randomly initialized weight vector; Step 2.4: According to the attention weight value calculated in step 2.3, the context vector y of each word i Perform weighted summation to obtain the text vector s of the attention mechanism layer: Step 2.5: The text word matrix obtained in step 1 is used as the input of the convolutional neural network. Assume that the size of the convolution kernel is w∈R h×d , where h represents the window width of the filter. When the convolution kernel is applied to a window consisting of h words, a new feature c can be generated. ij , the calculation formula is as follows: c ij =f(wX j:j+h-1 +b) In the formula, c ij represents the jth value on the i-th feature map; i is the number of convolution kernels, j is in the range of [1:n-h+1], indicating the number of times the convolution kernel can slide when the step size is 1; f is the ReLU activation function, and b is the bias term; Step 2.6: When the number of convolution kernels is m, all convolution kernels are convolved with the word vector matrix and the text word matrix to obtain m corresponding feature maps c = {c1, c2, ..., c n }; The pooling layer performs convolution on each feature map c i =[c i,1 ,c i,2 ,…,c i,n-h+1 ]Perform the maximum pooling operation and select the maximum value As the most important feature in the feature map, the other feature maps are similar, and the m eigenvalues obtained are concatenated After passing through the fully connected layer, the text vector d is obtained; Step 2.7: Concatenate the output s of the attention mechanism layer in step 2.4 and the output d of the convolutional neural network layer in step 2.6, pass them to the fully connected layer, and obtain the text feature vector z; Step 2.8: Calculate the cosine similarity based on the text feature vector and calculate the text feature vector z i and z j The cosine similarity sim ij : Where ||·|| is the L2 norm of the vector; Step 2.9: Construct a text network based on the similarity between text feature vectors, use text as nodes in the network, set a threshold θ, and if the similarity between two texts is sim ij >θ, then add an edge between the two texts; In step 3, K-means clustering is used to discover communities; the nodes in the network are clustered to explore the potential relationships between the nodes, and the following are obtained: Initial community center u=[u1,…,u j ,…,,u k ]; Community division result C = [C1,…,C j ,…,C k ].
6. The community discovery method according to claim 1, characterized in that: In step 4, a clustering loss function for the community discovery task is constructed based on text features and community centers, so that the community discovery task can guide feature learning. Finally, the Adam optimization function and the back propagation algorithm are combined to minimize the loss function, and the network model parameters are updated to improve the text features and community centers to obtain the final community division results.
7. The community discovery method according to claim 6, characterized in that: The step 4 comprises: Step 4.1: Construct clustering loss function LS: in, In the formula, q i =[q i1 ,...,q ij ,...,q ik ], q ij To use the Student's t-distribution function as the kernel to measure the sample point z i To the Community Center j The similarity of is the clustering soft assignment function; p i =[p i1 ,...,p ij ,...,p ik ], p ij Assign a function to the auxiliary target as the community discovery result q ij The true label of Step 4.2: Combine the Adam optimization function and the back propagation algorithm to minimize the clustering loss function, continuously reduce the difference between the clustering soft allocation function and the target allocation function to learn the results with higher confidence in community division, and update the network model parameters to optimize the text features and community centers. After network training, the community discovery result C i is the clustering soft assignment function q ij With the maximum probability, the text feature z i Assigned to community center j The label, that is 8. A community discovery system based on text deep clustering, characterized by: It includes a processor and a storage medium, wherein the storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, a community discovery method based on text deep clustering as described in any one of claims 1 to 7 is implemented.
9. A storage medium, the storage medium being a computer-readable storage medium for storing software program codes, characterized in that: The software program code is used to execute the community discovery method based on text deep clustering as described in any one of claims 1-7.
Citation Information
Patent Citations
Community discovery algorithm based on improved density peak clustering
CN109299749A
Cited By
Hot event dynamic monitoring method and device based on public information
CN120996962A