A short text stream clustering method based on episodic memory

By introducing plot memory mechanism and dynamic similarity threshold update in short text stream clustering, combined with online and offline clustering algorithms, the problems of sparseness and real-time in short text stream clustering are solved, and efficient short text stream clustering is achieved.

CN116166799BActive Publication Date: 2025-05-27CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211573825.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-08
Publication Date
2025-05-27
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

Existing short text stream clustering algorithms have limitations in dealing with the sparseness and real-time requirements of short texts, especially the difficulty in dealing with text sparseness and lack of real-time.

Method used

A short text stream clustering method based on plot memory is proposed. Through online clustering and plot memory mechanisms, the similarity threshold is dynamically calculated, clustering features are updated, and the algorithm performance is improved through cluster enhancement, semantic redistribution and deletion of outdated clustering algorithms in the offline stage.

Benefits of technology

This method can effectively deal with the sparseness problem of short text, improve real-time and clustering performance, reduce computing overhead, and ensure the feasibility of the stream clustering scheme.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116166799B_ABST
    Figure CN116166799B_ABST
Patent Text Reader

Abstract

The present invention discloses a short text stream clustering method based on episodic memory, which relates to the field of text data mining. Short text streams have characteristics such as infinite length, topic evolution, and sparse text features. Most of the existing short text stream clustering methods based on similarity require manual setting of a similarity threshold and cannot handle the text sparsity problem well. The present invention proposes a short text stream clustering method based on similarity of episodic memory. First, the idea of episodic memory is incorporated into the stream clustering algorithm. Then, the feature representation of clustering is enhanced through sparse experience replay, and the reverse index is used to improve the clustering efficiency. In the online stage, through a new similarity calculation formula and dynamic calculation of the similarity threshold, the current text is assigned to an existing cluster or a new cluster, and the clustering features are continuously updated. In the offline stage, the overall algorithm performance is improved through a clustering enhancement algorithm, a semantic reallocation algorithm, and an algorithm for deleting outdated clusters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data mining, and particularly relates to a short text stream clustering method with episodic memory based on similarity. Background Art

[0002] A short text stream refers to short text data that arrives continuously over time. Every day, more and more short text data is generated on social platforms and other websites, such as Weibo, Twitter, news websites, etc. Short text stream clustering has received increasing attention in recent years due to its diverse applications, such as hot topic detection, event detection and tracking, news recommendation, etc.

[0003] Different from traditional clustering algorithms (such as k-means, Gaussian mixture model, spectral clustering), short text stream clustering algorithms can be divided into similarity-based stream clustering algorithms and model-based stream clustering algorithms according to the technologies adopted. Most similarity-based short text stream clustering algorithms use the vector space model to represent documents, calculate the similarity between documents and clusters through similarity metrics, and then assign documents to existing clusters or new clusters according to similarity thresholds. Most of these algorithms run relatively fast and have good real-time performance. The limitations are that they need to manually set thresholds to determine the topic assignment of online documents and are difficult to handle the problem of text sparsity. Model-based short text stream clustering algorithms assume that documents are generated by a mixture model, which selects topics with a certain probability and then selects words with a certain probability from these topics to generate documents. Usually, Gibbs sampling or EM algorithms are used to estimate the parameters of the mixture model. The disadvantages are that the number of topics needs to be specified in advance and they cannot handle the evolving unknown number of topics in short text streams. In addition, there are mainly two text scanning methods in short text stream clustering: batch processing-based and one-pass-based methods. The batch processing-based method performs clustering on the text in each batch multiple times, and the real-time performance is not good when dealing with large-scale text streams. The one-pass-based method processes each text only once, with high operating efficiency but unable to handle the sparsity problem of short texts well. Therefore, the present invention designs a short text stream clustering method with episodic memory. Summary of the Invention

[0004] In order to overcome the deficiencies of the above-mentioned prior art, a short text stream clustering method based on episodic memory is proposed. The technical solution of the present invention is as follows:

[0005] A short text stream clustering method based on episodic memory, comprising the following steps:

[0006] Online clustering, by using a similarity calculation formula and dynamically calculating a similarity threshold, allocating the current text to an existing cluster or a new cluster, and continuously updating the clustering features;

[0007] Episodic memory, by selecting texts from the episodic memory module at regular time intervals for experience replay, helps to update the feature vectors of these topics using past text information, increasing the probability of these topics being selected during the clustering process;

[0008] Offline clustering improves the overall algorithm performance through clustering enhancement algorithms, semantic reassignment algorithms, and deletion of outdated clustering algorithms.

[0009] Furthermore, the online clustering uses a single pass. First, the current text is preprocessed and feature extracted. If the number of processed texts reaches the experience replay interval R, E texts are randomly selected for sparse experience replay to update the clustering features, and then the current text is clustered. According to the reverse index, the existing cluster containing the text features is selected, and the current text is assigned to the existing cluster or a new cluster using a similarity calculation formula, and the clustering features are continuously updated.

[0010] More specifically, the feature extraction includes text feature extraction from both the lexical feature and semantic levels. The lexical feature extracts text features at the lexical level through biterm. Biterm tokenizes the preprocessed text, and then calculates the Cartesian product of the word list. Biterm uses the following formula to implement feature extraction:

[0011] f(k) = {{w i , w j}, i, j ∈ [1, k], i ≠ j}

[0012] f(k) represents the feature extraction of text t, where k is the number of words in the text, w i and w j are different words in the text. Then, the document vector is constructed by the word averaging method to represent the text semantic information. The word vector is obtained through the Glove model. The lexical feature of each cluster is represented by a CF vector:

[0013]

[0014] where is the frequency of feature f in cluster z, n z is the number of features in cluster z, m z is the number of texts in cluster z, id z is the unique id of cluster z. The semantic representation of each cluster consists of the cluster vector S z and the cluster center vector composed. S z is the sum of the document vectors of the texts in cluster z, and is calculated by dividing the cluster vector by the cluster size.

[0015] Further, the similarity calculation formula uses a similarity calculation formula based on common features to calculate the similarity between the text and the cluster. The similarity calculation formula is as follows:

[0016]

[0017] First, sum the frequencies corresponding to the common feature f between the text t and the cluster z, and then multiply by a weight FI f , and the calculation formula is as follows:

[0018]

[0019] Furthermore, the updated cluster feature means that when the text t is assigned to the cluster z, the CF vector of the cluster z is updated. The specific update steps are as follows:

[0020]

[0021] n z = n z + N t

[0022] m z = m z + 1

[0023] S z = S z + S t

[0024]

[0025] Where is the frequency corresponding to the feature f in the cluster z, is the frequency corresponding to the feature f in the replay text t, n z is the number of features of the cluster z, N t is the number of features of the text t, m z is the number of texts in the cluster z, S z is the cluster vector of the cluster z, is the cluster center vector of the cluster z;

[0026] For the update processing of the cluster id, it is as follows:

[0027]

[0028] If the current text is not assigned to a new cluster, then id z remains unchanged, otherwise id z increases by 1. Therefore, the most recently created cluster has the highest cluster id. At the same time, update the reverse index, and add the cluster id to the corresponding CF feature vector for each feature in the text.

[0029] When the number of continuously arriving documents reaches the storage interval, the current text is added to the episodic memory module. During the clustering process, texts are selected from the episodic memory module at certain time intervals for experience replay. In the online clustering stage, a forward index of id-clustering features is generated in memory, and then a reverse index is created through the forward index. Through the reverse index, the clustering ids including the same features can be found, and clustering is performed by calculating the similarity between the current text and the selected clusters with common features.

[0030] Specifically, the clustering enhancement algorithm selects a group of clusters obtained by the online clustering module at each update interval to enhance the distribution of these clusters. The size of the cluster corresponds to the number of texts in the cluster. Clusters with a size greater than μ + σ are selected, where μ and σ are the mean and variance of the sizes of all clusters in the online clustering results respectively. Clustering enhancement is performed through iterative classification. In each iteration, training sets and test sets containing non-outliers and outliers are generated respectively. The classification algorithm is trained using the training set, and then the trained model is used to classify the test set. The iteration is repeated until the text distribution in each cluster tends to be stable or the preset maximum number of iterations is reached.

[0031] Specifically, the semantic reallocation algorithm reallocates single-text clusters. For the set of texts T in a cluster with a size of 1, the texts t in it are first preprocessed to obtain a word list W t , and then the text semantic feature vectors are accumulated with SUM. Finally, through the word averaging method

[0032]

[0033] Specifically, the semantic vector of the text is obtained, and the similarity Sim t between the semantic vector of the text and the cluster center vector of the existing cluster is calculated. Then, the maximum similarity max t in Sim t , the mean similarity μ t and the variance σ t are calculated. If max t > μ t + σ t t , then the clustering label of the current text is modified to the clustering label j corresponding to max τ . Otherwise, the text t remains in the original cluster.

[0034] For the obsolete cluster deletion algorithm, since the clusters recently created by the online clustering algorithm have higher ids z , the means μ z , μ m of the cluster numbers and cluster sizes are calculated respectively, and the variance σ mz , σ m , delete the cluster number id z less than μ z -σ z and the cluster size is less than μ m -σ m the CF vectors corresponding to the clusters and the information of the clusters in the reverse index F.

[0035] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the short text stream clustering method based on episodic memory as described above.

[0036] The advantages and beneficial effects of the present invention are as follows:

[0037] The efficient short text stream clustering method based on episodic memory combines episodic memory and online-offline clustering assignment. First, the idea of episodic memory is incorporated into the stream clustering algorithm, then the feature representation of the clustering is enhanced through sparse experience replay, and the reverse index is used to improve the clustering efficiency; in the online stage, the current text is assigned to the existing cluster or a new cluster through a new similarity calculation formula and dynamically calculating the similarity threshold, and the clustering features are continuously updated; in the offline stage, the overall algorithm performance is improved through the clustering enhancement algorithm, the semantic reallocation algorithm, and the algorithm for deleting outdated clusters. In addition, the increased computational overhead in the clustering process is also very limited, which can ensure the feasibility of the stream clustering scheme of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a model diagram of a short text stream clustering method based on episodic memory in an embodiment of the present invention;

[0039] Figure 2 is a forward index diagram of cluster id-cluster feature generated in the online clustering stage of the present invention;

[0040] Figure 3 is a reverse index diagram of cluster feature-cluster id generated in the online clustering stage of the present invention;

[0041] Figure 4 is a structural diagram of the episodic memory module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0043] Reference Figure 1 , Figure 2 , Figure 3 and Figure 4 , the specific implementation manner of the present invention is as follows:

[0044] 1. The online clustering stage is based on a one-pass method. First, preprocess the current text (including word segmentation and stop word removal) and extract features. If the number of processed texts reaches the experience replay interval R, randomly select E texts for sparse experience replay to update the clustering features, and then cluster the current text. Select the existing clusters containing the text features according to the reverse index, and use a custom similarity calculation formula to assign the current text to the existing cluster or a new cluster, and continuously update the clustering features. Specifically, it includes short text feature extraction, similarity calculation, and construction of a clustering model. The text to be processed is extracted from two levels: lexical features and semantics. The lexical features are extracted from the text at the lexical level through biterm. Biterm performs word segmentation on the preprocessed text, and then calculates the Cartesian product of the word list. The number of extracted features is more than that of other methods, and it can more comprehensively extract the lexical features in the text. Biterm uses the following formula:

[0045] f(k) = {{w i , w j}, i, j ∈ [1, k], i ≠ j}

[0046] to extract the text features. The number of extracted features is (k × (k - 1)) / 2. f(k) represents the feature extraction of text t, where k is the number of words in the text, and w i and w j are different words in the text. Then, the document vector is constructed by the word averaging method to represent the semantic information of the text, where the word vector can be obtained through the Glove model. The lexical features of each cluster are represented by the CF vector:

[0047]

[0048] where is the frequency corresponding to the feature f in cluster z, n z is the number of features in cluster z, m z is the number of texts in cluster z, and id z is the unique id of cluster z. The semantic representation of each cluster consists of the cluster vector S z and the cluster center vector . S z is the sum of the document vectors of the texts in cluster z. is calculated by dividing the cluster vector by the cluster size (represented by the number of texts m z in the cluster).

[0049] In the online clustering stage, a similarity calculation formula based on common features is used to calculate the similarity between the text and the clusters. First, a forward index of id-cluster features is generated in memory, and then a reverse index of cluster feature-id is created through the forward index, as Figure 2 shown. Through the reverse index, the cluster ids including the same feature can be found. Clustering is performed by calculating the similarity between the current text and the selected clusters with common features. Using the reverse index can reduce the number of text similarity calculations and improve the running efficiency of the algorithm. The specific similarity calculation formula is as follows:

[0050]

[0051] First, sum the frequencies corresponding to the common feature f between the text t and the cluster z, and then multiply by a weight FI similar to the inverse document frequency (IDF). f , and the calculation formula is as follows:

[0052]

[0053] where |z| is the total number of existing clusters, and |f∈z| is the number of clusters containing the feature f. The size of FI f can reflect the importance of the feature f. If the feature f appears in more clusters, the feature is less important and the weight is lower, and vice versa.

[0054] After obtaining the cluster assignment of the current text in the online clustering stage, a cluster model needs to be constructed. Specifically, when the text t is assigned to the cluster z, the CF vector of the cluster z is updated. The specific update steps are as follows:

[0055]

[0056] n z = n z + N t

[0057] m z = m z + 1

[0058] S z = S z + S t

[0059]

[0060] where is the frequency corresponding to the feature f in the cluster z, is the frequency corresponding to the feature f in the text t, n z is the number of features in the cluster z, and N tis the number of features of text t, m z is the number of texts in cluster z. S t is the semantic vector of text t, S z is the cluster vector of cluster z, is the cluster center vector of cluster z. For the update of the cluster id, the processing in the online clustering stage is as follows:

[0061]

[0062] If the current text is not assigned to a new cluster, then id z remains unchanged, otherwise id z increases by 1. Therefore, the most recently created cluster has the highest cluster id. Subsequently, the obsolete cluster deletion algorithm can be operated according to the cluster id, and then the inverted index is updated where is the set of cluster ids of the clusters that include feature f in the cluster features, and add the cluster record id corresponding to the current text feature. Finally, add the cluster id to the corresponding CF feature vector for each feature in the text, and the specific operation is as in the above update steps.

[0063] The scenario memory module is incorporated into the online clustering stage, as shown in the appendix Figure 4 It is used to store some previously processed texts in memory. Due to limited memory, the present invention selects to write new texts into memory at a certain time interval R. When the number of continuously arriving documents reaches the storage interval, the current text is added to the episodic memory module. When the replay interval is reached, randomly select E texts in the episodic memory module and perform a scan clustering on them. For these previously processed texts, only select an existing cluster instead of generating a new cluster. After determining the cluster through sparse experience replay, it is necessary to update the vocabulary features and semantic representations corresponding to the cluster. The specific update steps are the same as the steps for constructing the cluster model.

[0064] After the online clustering stage, the clustering information of the current text can be basically determined, but we can enhance the effect of online clustering through offline clustering. The offline clustering stage includes an enhanced clustering algorithm, a semantic reallocation algorithm, and an obsolete cluster deletion algorithm.

[0065] The enhanced clustering algorithm selects a group of clusters obtained by the online clustering module at each update interval and enhances the distribution of these clusters. The size of the cluster corresponds to the number of texts in the cluster. Select the clusters with a cluster size greater than μ z +σ z μ z and σ zThey are the mean and variance of the sizes of all clusters in the online clustering results respectively. Clusters with a large number of texts have more outliers, which may lead to incorrect assignment of future texts. By enhancing the text distribution in larger clusters, future texts are assigned to appropriate clusters. This algorithm performs cluster enhancement through iterative classification. In each iteration, training sets and test sets containing non-outliers and outliers are generated respectively. The classification algorithm is trained using the training set, and then the trained model is used to classify the test set. The iteration is repeated until the text distribution in each cluster tends to be stable or the preset maximum number of iterations is reached.

[0066] The semantic redistribution algorithm in the offline clustering stage redistributes single-text clusters (clusters containing only one text). The semantic similarity between the text and the cluster center is calculated by cosine similarity, and then the mean value μ t and variance σ t of the similarities are calculated respectively. If the maximum similarity value is greater than μ t +σ t , the text is assigned to the corresponding cluster, otherwise it remains in the cluster obtained through the online clustering stage. Specifically, for the set of texts T in a cluster with a size of 1, the text t in it is first preprocessed to obtain a list of words W t , then the semantic feature vectors of the text and SUM are accumulated, and finally the semantic vector of the text is obtained by the word averaging method

[0067]

[0068] . The similarity Sim t between S calculated by cosine similarity and the cluster center vector of the existing cluster t is calculated, and then the maximum similarity max t in Sim t , the similarity mean μ t and variance σ t are calculated. If max t >μ t +σ t , the current text cluster label is modified to the cluster label j corresponding to max t , otherwise the text t remains in the original cluster. After semantic assignment, the lexical feature representation and semantic representation of the cluster need to be updated simultaneously.

[0069] The obsolete cluster deletion algorithm in the offline clustering stage deletes obsolete clusters according to the cluster number id z and the cluster size. According to the online clustering algorithm, the recently created clusters have higher ids z . The mean values μ z , μ mSum of variances σ z , σ m , delete the cluster number id z less than μ z -σ z and the cluster size is less than μ m -σ m the CF vectors corresponding to the clusters and the information of the clusters in the reverse index F.

Claims

1. A short text stream clustering method based on episodic memory, characterized in that, it includes the following steps: Online clustering, by using a similarity calculation formula and dynamically calculating a similarity threshold, allocating the current text to an existing cluster or a new cluster, and continuously updating the clustering features; The online clustering adopts a single pass. First, preprocess and extract features from the current text. The feature extraction includes extracting text features from two levels: lexical features and semantics. The lexical features are extracted from the text at the lexical level through biterm. Biterm tokenizes the preprocessed text, and then calculates the Cartesian product of the word list. The biterm uses the following formula to implement feature extraction: f(t) = {{w i , w j}, i, j ∈ [1, k], i ≠ j} f(k) represents the feature extraction of text t, where k is the number of words in the text, and w i and w j are different words in the text. Then, the document vector is constructed by the word averaging method to represent the text semantic information. The word vectors are obtained through the Glove model, and the lexical features of each cluster are represented by a CF vector: Among them is the frequency corresponding to the feature f in the clustering z , n z is the number of features of the clustering z , m z is the number of texts of the clustering z , id z is the unique id of the clustering z . The semantic representation of each clustering consists of the clustering vector S z and the clustering center vector . S z is the sum of the document vectors of the texts in the clustering z is calculated by dividing the clustering vector by the clustering size; If the number of processed texts reaches the experience replay interval R, randomly select E texts for sparse experience replay to update the clustering features, then cluster the current text. Select the existing clusters containing the text features according to the reverse index, and use the similarity calculation formula to allocate the current text to an existing cluster or a new cluster, and continuously update the clustering features; The update of the clustering features means updating the CF vector of cluster z when text t is allocated to cluster z. The specific update steps are as follows: n z = n z + N t m z = m z + 1 S z = S z + S t where is the frequency corresponding to the feature f in the cluster z, is the frequency corresponding to the feature f in the replay text t, n z is the number of features of the cluster z, N t is the number of features of the text t, m z is the number of texts of the cluster z, S z is the cluster z 's cluster vector, is the cluster z 's cluster center vector, S t is the semantic vector of the text t; The update processing for the cluster id is as follows: If the current text has not been assigned to a new cluster, then the id z remains unchanged, otherwise the id z is incremented by 1. Thus, the most recently created cluster has the highest cluster id, and the inverted index is updated by adding the cluster id to the corresponding CF feature vector for each feature in the text; Episodic memory, select texts from the episodic memory module for experience replay every certain time interval; Offline clustering, improve the overall algorithm performance through a clustering enhancement algorithm, a semantic reassignment algorithm, and a deletion of obsolete clusters algorithm; The clustering enhancement algorithm selects a group of clusters obtained by the online clustering module at each update interval to enhance the distribution of these clusters. The size of the cluster corresponds to the number of texts in the cluster. Select the clusters with a size greater than μ + σ, where μ and σ are the average and variance of the sizes of all clusters in the online clustering results respectively. Perform clustering enhancement through iterative classification. Each iteration generates a training set and a test set containing non-outliers and outliers respectively. Use the training set to train the classification algorithm, and then use the trained model to classify the test set. Repeat the iteration until the text distribution in each cluster tends to be stable or reaches the preset maximum number of iterations; The semantic redistribution algorithm redistributes single-text clustering. For the clustering text set T with a clustering size of 1, first preprocess the text t therein to obtain a word list W t , then perform text semantic feature vector and SUM accumulation, and finally use the word averaging method Obtain the semantic vector of the text, and calculate S through cosine similarity t and the cluster center vector of the existing cluster to get the similarity Sim t , then calculate the maximum similarity max t in Sim t , the mean similarity μ t and the variance σ t . If max t >μ t +σ t , then modify the current text cluster label to the cluster label j corresponding to max t , otherwise the text t remains in the original cluster.

2. The short text stream clustering method based on episodic memory according to claim 1, characterized in that: The similarity calculation formula uses a similarity calculation formula based on common features to calculate the similarity between the text and the cluster. The similarity calculation formula is as follows: First, sum the frequencies corresponding to the common feature f between the text t and the cluster z , and then multiply by a weight FI f . The calculation formula is as follows:

3. The short text stream clustering method based on episodic memory according to claim 1, characterized in that: When the number of continuously arriving documents reaches the storage interval, add the current text to the episodic memory module. During the clustering process, select texts from the episodic memory module for experience replay every certain time interval. In the online clustering stage, a forward index of id-clustering features will be generated in memory, and then a reverse index will be created through the forward index. Through the reverse index, the cluster id including the same feature can be found, and clustering is performed by calculating the similarity between the current text and the selected clusters with common features.

4. The short text stream clustering method based on episodic memory according to claim 1, characterized in that: The described deletion of obsolete clustering algorithm is based on the fact that the clusters created most recently by the online clustering algorithm have higher IDs z , and calculate the mean values μ z , μ m and variance σ z , σ m , delete the cluster ID id z less than μ z - σ z and the cluster size less than μ m - σ m of the CF vectors corresponding to the clusters and the information of these clusters in the reverse index F.

5. A computer-readable storage medium, characterized in that: a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the method for clustering short text streams based on episodic memory according to any one of claims 1 to 4.