A method for constructing a context-enhanced Dirichlet model that supports online clustering of short text streams
By constructing the context-enhanced Dirichrey model, the efficiency and robustness of short text flow clustering in high-speed and conceptual drift environments are solved, and efficient online clustering is achieved, reducing sparseness and improving the performance of the model.
Patent Information
- Application Number
- CN202211504585.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-29
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-11-29
AI Technical Summary
Existing short text stream clustering methods are inefficient and ineffective in processing high-speed, conceptual drift, and time-limited data, making it difficult to effectively handle ever-changing term weights and microclustered streaming documents.
A context-enhanced Dirichrey model is constructed, adding or creating new clusters by calculating probability selection, merging similar clusters, deleting obsolete clusters, utilizing unique neighbor group distribution and situational reasoning processes to reduce sparsity, and using window-based term co-occurrence matrix and word-specific weight calculations.
The efficiency and robustness of online clustering of short text streams is improved, cluster sparsity is reduced, and model performance in NMI, homogeneity and cluster purity is improved.
Smart Images

Figure CN115827861B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and natural language processing, and particularly relates to a method for constructing a context-enhanced Dirichlet model that supports online clustering of short text streams. Background Art
[0002] In the past decade, a large amount of short text data has been generated on social media every day, such as tweets, Facebook posts, and Q&A platforms. In recent years, the clustering of such continuously arriving short texts has received extensive attention due to its various applications in topic tracking, news recommendation, rumor detection, etc.
[0003] The text representation of documents is an important task in natural language processing. The representations usually obtained are based on selecting a suitable set of terms (low-dimensional vocabulary) and their relative importance (weighting scheme) to capture the document content. Specifically, the term weighting scheme helps to capture the context of words in a specific sentence or document. For example, the bag-of-words (BoW) representation and the TF-IDF weighting scheme are widely used as term indexing and importance scores respectively. A method for learning term specificity (weighting scheme) has been proposed for generating word embeddings in a static corpus. However, in a streaming environment, such a representation cannot be directly mapped. Especially for the evolving term weights, and for streaming documents with micro-clusters that contain overlapping term subspaces, the term distribution is unknown and may change over time.
[0004] Different from long text documents (such as blogs and academic papers), since short texts contain few words, processing short texts is crucial and challenging. However, due to the unique properties of the stream (i.e., high speed, concept drift, and evolution), the clustering task becomes more complex because it needs to process data with time and space limitations. Different from static text documents, the occurrence of concept evolution and concept drift is unknown in high-speed text streams, thus generating the need to process the stream in an online manner. Summary of the Invention
[0005] Aiming at the deficiencies in the prior art, the present invention provides a method for constructing a context-enhanced Dirichlet model that supports online clustering of short text streams, which has the advantages of high model efficiency and strong robustness.
[0006] The above technical object of the present invention is achieved through the following technical solutions:
[0007] A method for constructing a context-enhanced Dirichlet model that supports online clustering of short text streams, comprising the following steps:
[0008] Step 1, according to the calculated probability, select to add the arriving document to the active clusters of the model, or create a new cluster for addition;
[0009] Step 2, when the probability that the documents of an existing cluster in the model arrive is less than the pseudo-probability, the document is regarded as the emergence of a new topic, thereby creating a new cluster;
[0010] Step 3, as the documents arrive, the model checks and deletes the old clusters (i.e., outdated topics), so that the recent topic clusters of the current distribution are active in the model. To infer the number of active clusters in the model, the random number of documents η from the most recent ψ documents is resampled after each ρ time unit interval.
[0011] 3. The present invention is further configured such that each micro-cluster in the model is represented as a 6-tuple where z represents the micro-cluster, mz is the number of documents therein, is a two-dimensional matrix containing the terms and their frequencies in the z documents, N z is the total number of words in the z documents, is the total number of words in the d document, and the tuple element cw z is a three-dimensional matrix for storing the co-occurrence of terms in ; w i and w j The score between them is defined as
[0012]
[0013] where, is the word frequency of w i in the document d, l z and u_{z} store the decay weight and the last update timestamp of the cluster. The CF set has two important attributes, including addable and removable. These attributes allow the cluster to be updated incrementally over time. The Addable attribute enables the cluster to be updated by adding new documents to it and is defined.
[0014] The present invention is further configured as: Definition 1, by using the addable attribute to update the cluster, the document d can be added to the cluster z:
[0015] m z = m z + 1
[0016]
[0017] cw z = cw z ∪ cw d
[0018] N z = N z + N d ;
[0019] Definition 2. Deleting document d from cluster z using the deletable attribute:
[0020] m z = m z - 1
[0021]
[0022] cw z = cw z - cw d
[0023] N z = N z - N d ,
[0024] where N d represents the total number of words in the document, and cw d is the document co - occurrence matrix based on the window size. The co - occurrence matrix cw of the document d contains the frequency ratio between two adjacent terms, defined as
[0025]
[0026] where is the word frequency of w i in the document with (w i , w j ) ∈ d as the theme.
[0027] The present invention is further set as follows: Initially, in the model, there is no cluster. Therefore, a new empty cluster is created and the first - arriving document is added to it. The next - arriving document should be added to the existing (active) cluster z, (z ∈ M) of the model, or it will result in the creation of a new cluster; calculating the similarity between document d and the active cluster z as,
[0028]
[0029] The present invention is further set as follows: A weight is proposed to calculate the specificity of a term. If a term co - occurs with different terms each time, it indicates that the term is not very specific and can thus be used with multiple concepts. However, a highly specific term has a smaller number of co - occurring neighboring data. In addition, the normal neighboring data of a term depends on the defined window size. Therefore, the word specificity S of w i is defined as:
[0030]
[0031] where δ is the adjacent window size, and g(w i ) is defined as
[0032]
[0033] g(w i ) is used to calculate the ratio between the neighborhood population and the defined window size (σ) and its frequency in the document, where, is the unique neighboring data of the term w i in the model, is the total term frequency, and the fraction is calculated by using the boundary of the window size defined by the hyperbolic tangent sigmoid function.
[0034] The present invention is further configured such that, during initialization, a new cluster is created for the first document of the stream, and in order to automatically identify new topics in the arriving documents over time, by transforming the formula p(z d |G z ) = p(G z ).p(d|G z ), the probability p(z new |d) is derived as follows,
[0035]
[0036] where, αD represents the pseudo-population of the document, and V z is the average vocabulary of the active clusters The value of β helps to calculate the pseudo-word similarity with the new clusters; the equation gives the condition for creating a new cluster.
[0037] The present invention is further configured such that a window-based term co-occurrence matrix is introduced, where each term can only be paired with neighboring data having distance while maintaining the sentence order.
[0038] The present invention is further configured such that, in order to maintain the current distribution of concepts, the model needs to maintain the active clusters and remove the outdated clusters, and the decay weight of each cluster is updated over time as
[0039]
[0040] where, t c represents the current timestamp of the model, and the last update timestamp of the cluster is stored. Initially, the decay weight of each new cluster is set to 1. If l_{z} is approximately zero, the cluster is processed for deletion from the model, that is, the micro-cluster cannot capture the current entry distribution of the topics in the text stream.
[0041] The beneficial effects of the present invention are as follows: A context-enhanced Dirichlet model that supports online clustering of short text streams. This solution utilizes the distribution of the unique neighbor groups of the entire model. For active clusters, a scenario inference process is proposed, which reduces the cluster sparsity of the model. In addition, EINDM also merges highly similar clusters, automatically generating cluster populations that are close to the actual clusters. We conducted extensive empirical analyses to demonstrate the efficiency and robustness of the proposed model while observing the results within different parameter ranges. EINDM generates a new word-specific term weight scheme using the distribution of the unique neighbor groups of the entire model. At the same time, the scenario inference process reduces the cluster sparsity of the model. EINDM also merges highly similar clusters, automatically generating cluster populations that are close to the actual clusters. Compared with the most recent state-of-the-art clustering models, EINDM has the best performance in terms of NMI, homogeneity, and clustering purity. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of the algorithm flow;
[0043] Figure 2 It is a schematic diagram of an example of a word window, where the crossed-out part represents a non-movable term window. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The technical solutions in the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0045] Model the topic generation of the Dirichlet distribution, where the most important task is to formulate the relationship between the arriving documents and the topics (micro-clusters) according to the probability distribution.
[0046] Step 1: According to the calculated probability, each arriving document is either added to the active clusters (active topics) of the model or a new cluster is created (a new topic has arrived).
[0047] Step 2: If the probability of the arrival of a document that has already been clustered in the model is less than the pseudo-probability, the document is regarded as the emergence of a new topic, thereby creating a new cluster.
[0048] Step 3: As documents arrive, the model checks whether the old clusters (obsolete topics) are deleted.
[0049] In this way, the cluster of recent topics with the current distribution is active in the model. To infer the number of active clusters in the model, the number of random documents η from the most recent ψ documents is resampled after each ρ time unit interval. Each key process is described in detail below.
[0050] 1. Future set of clusters. Each micro-cluster in the model is represented as a 6-tuple where z represents the micro-cluster, mz is the number of documents therein, is a two-dimensional matrix that contains the terms in document z and their frequencies, N z is the total number of words in document z, is the total number of words in document d, the tuple element cw z is a three-dimensional matrix for storing the co-occurrence of terms in ; w i and w j The score between is defined as,
[0051]
[0052] where, is the word frequency of w in document d i . l z and u_{z} store the decay weight and the timestamp of the last update of the cluster. The CF set has two important properties: (i) addable and (ii) removable. These properties allow the cluster to be updated incrementally over time. The Addable property enables the cluster to be updated by adding new documents to it and is defined as, Definition 1: By using the addable property to update the cluster, document d can be added to cluster z.
[0053]
[0054] Definition 2: Use the removable property to remove document d from cluster z.
[0055]
[0056] Here, N d represents the total number of words in the document. cw d is the document co-occurrence matrix based on the window size. The co-occurrence matrix cw of the document d contains the frequency ratio between two adjacent terms and is defined as:
[0057]
[0058] 2. Document Cluster Similarity: Initially, there are no clusters in the model, so a new empty cluster is created and the first arriving document is added to it. The next arriving document should be added to the existing (active) cluster z (z ∈ M) in the model (based on Equation 7), or it will result in the creation of a new cluster (based on Equation 13). To calculate the similarity between document d and the active cluster z} (the model is represented as ),
[0059]
[0060] The definitions of all symbols are shown in the following table.
[0061]
[0062] Table 1: Symbols and Marks
[0063] Here, the first part of this equation shows the cluster popularity p(G_{z}), while the remaining two parts are for calculating the homogeneity p(d|G_{z}). The second part of the equation is responsible for capturing the similarity in the single-item space. The third part solves the term ambiguity problem by calculating the similarity in the semantic term space (i.e., the co-occurrence of term sets). The second part of the homogeneity, where the value of β serves as the pseudo-weight for unseen words, is based on the multinomial distribution (i.e., ). Using the term occurrences in the cluster while capturing the homogeneity (similarity). However, different from the static environment where the inverse document frequency can be used to calculate the term importance in the global space, the term-document distribution is unknown. Therefore, we define a similar weighted score, called the inverse cluster frequency ICF to calculate the importance of a term, defined as: w The denominator part of Equation 8 is the total number of active clusters in the model that contain the word w, and the nominator is the total number of active clusters in the model. This means that more clusters containing the word w are less important than the word. If the word w is contained in few clusters, it indicates a higher importance. To capture this behavior, a new word-specific S
[0064]
[0065] weight is introduced, which is defined in Equation 9. w 3. Word Specificity: A new weight is proposed to calculate the specificity of a term. The basic idea is that if a term co-occurs with different terms each time, it indicates that the term is less specific and thus can be used with multiple concepts. However, highly specific terms have a smaller number of co-occurring neighbors. In addition, the normal neighbors of a term depend on the defined window size (as a model parameter). Therefore, considering the co-occurrence window size constraint, we define the word specificity S of w
[0066] as, i Here, δ is the adjacent window size, and g(w
[0067]
[0068] ) is defined as, i as,
[0069]
[0070] g(w i ) calculates the ratio between the neighborhood population and the defined window size (σ) and its frequency in the document. Here, is the number of unique neighbors of the term w in the model i , and is the total term frequency. The fraction is calculated by defining the bounds of the window size using the hyperbolic tangent sigmoid function.
[0071] 4. Automatic cluster creation: At initialization, the first document of the stream creates a new cluster. To automatically identify new topics in the arriving documents over time, we need a probability to help detect new distributions of words to create new clusters in case the document does not belong to any active cluster. By transforming the formula p(z d |G z ) = p(G z ).p(d|G z ), the probability p(z new |d) is derived as follows.
[0072]
[0073] Here, αD represents the pseudo-population of the document, and V z is the average vocabulary size of the active clusters The value of β helps calculate the pseudo-word similarity to the new cluster. Equation 14 gives the condition for creating a new cluster (see line 10 of Algorithm 1).
[0074]
[0075]
[0076] Algorithm 1
[0077] 5. Window-based co-occurrence matrix: This patent introduces a window-based term co-occurrence matrix where each term can only be paired with neighbors within a distance while maintaining the sentence order. An example is shown in Figure 2 . In this way, the co-occurrence matrix of a document cw d can have at most O(δN d ) entries. Here, N d is the document length. The weight between two adjacent terms w i and w j in the cluster is defined in Equation 6, where i ≠ j, (i - δ) ≤ j ≤ (i + δ) and δ ≥ 1
[0078] 6. Episodic Reasoning: The reasoning process has been proven to be applicable to the generation process to reduce cluster sparsity. Their iterative process has two main drawbacks: (i) A batch may not capture the current distribution, and (ii) it increases the processing time cost, making it unsuitable for high-speed streams. In contrast, we propose an episodic reasoning procedure that not only effectively reduces the processing cost but also can cover the distribution of the stream.
[0079] 7. Deleting Outdated Clusters: To maintain the current distribution of concepts, the model needs to keep active clusters (current concepts) and remove outdated clusters (old concepts). A decay mechanism based on the flow velocity is adopted, which updates the importance score (decay weight) of clusters over time. If the model has not received documents recently, the decay weight (l z ) of each cluster in the model decreases over time. The decay weight of each cluster over time is updated as
[0080]
[0081] Here, t c represents the current timestamp of the model, and z stores the last update timestamp of the cluster. Initially, the decay weight of each new cluster is set to 1 (see line 10 of Algorithm 2). If l_{z} is approximately zero, the cluster is processed for deletion from the model (see line 6 of Algorithm 3), that is, the micro-cluster cannot capture the current term distribution of the topic in the text stream.
[0082]
[0083]
[0084] Algorithm 2
[0085] The present invention proposes a context-enhanced Dirichlet model for supporting online clustering of short text streams. This scheme utilizes the distribution of the unique neighbor groups of the entire model. For active clusters, an episodic reasoning process is proposed, which reduces the cluster sparsity of the model. In addition, EINDM also merges highly similar clusters, automatically generating a cluster population close to the actual clusters. We have conducted extensive empirical analyses to demonstrate the efficiency and robustness of the proposed model while observing the results within different parameter ranges. EINDM generates a new word-specific term weight scheme by utilizing the distribution of the unique neighbor groups of the entire model. At the same time, the episodic reasoning process reduces the cluster sparsity of the model. EINDM also merges highly similar clusters, automatically generating a cluster population close to the actual clusters. Compared with the most recent state-of-the-art clustering models, EINDM has the best performance in terms of NMI, homogeneity, and clustering purity.
[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for constructing a context-enhanced Dirichlet model to support online clustering of short text streams, characterized in that: Including the following steps: Step 1, according to the calculated probability, select to add the arriving document to the active cluster of the model, or create a new cluster for addition; Step 2, when the probability of the arrival of the documents in the existing clusters in the model is less than the pseudo-probability, the documents are regarded as the emergence of a new topic, thus creating a new cluster; Step 3, as the documents arrive, the model checks and deletes the old clusters, that is, the outdated topics, so that the recent topic clusters of the current distribution are active in the model. To infer the number of active clusters in the model, resample the random number η of documents from the most recent ψ documents after each ρ time unit interval; Initially in the model, there are no clusters, so a new empty cluster is created and the first arriving document is added to it. The next arriving document should be added to the existing active cluster z of the model, z ∈ M, or will result in the creation of a new cluster; At initialization, the first document of the text stream creates a new cluster. To automatically identify new topics in the arriving documents over time, by the transformation formula p(z d |G z ) = p(G z ).p(d|G z ), the probability p(z new |d) is derived as follows, Among them, is the word frequency matrix of w in document d, N d is the total number of words in document d, αD represents the pseudo-population of the document, V z is the average vocabulary of the active cluster The value of β helps calculate the pseudo-word similarity with the new cluster; the equation gives the conditions for creating a new cluster; A window-based term co-occurrence matrix is introduced, where each term can only be paired with neighboring data with distance while maintaining the sentence order; To maintain the current distribution of concepts, the model needs to maintain the active clusters and remove the outdated clusters. The decay weight of each cluster over time is updated to where t c represents the current timestamp of the model and stores the last update timestamp u z of the cluster. The decay weight of each new cluster is set to 1. If l z is approximately zero, the cluster is processed for deletion from the model, i.e., the cluster cannot capture the current term distribution of the topic in the text stream.
2. The method for constructing a context-enhanced Dirichlet model supporting online clustering of short text streams according to claim 1, wherein: Each micro-cluster in the model is represented as a 6-tuple where z represents the micro-cluster, m z is the number of documents in it, is a two-dimensional matrix containing the terms and their frequencies in the z documents, N z is the total number of words in the z documents, is the total number of words in the d documents, and the tuple element cw z is a three-dimensional matrix used to store the co-occurrence of terms in ; w i and w j The score between them is defined as Among them, is the word frequency of w in document d, i l z and u z store the decay weight and the timestamp of the last update of the cluster respectively. The CF set has two important attributes, including addable and removable. These attributes allow the cluster to be updated incrementally over time. The Addable attribute enables the cluster to be updated by adding new documents to it and is defined.
3. A method for constructing a context-enhanced Dirichlet model to support online clustering of short text streams according to claim 2, characterized in that: Definition 1, update the cluster by using the addable attribute, and add the document d to the cluster z: m z = m z + 1 cw z = cw z ∪ cw d N z = N z + N d ; Definition 2, use the deletable attribute to delete the document d from the cluster z: m z = m z -1 cw z = cw z - cw d N z = N z -N d , Among them, N d represents the total number of words in the document, cw d is the document co-occurrence matrix based on the window size, and the co-occurrence matrix cw of the document d contains the frequency ratio between two adjacent terms, defined as Among them, is w i The word frequency with w i , w j ) ∈ d as the theme.
4. The method for constructing a context-enhanced Dirichlet model that supports online clustering of short text streams according to claim 3, characterized in that: Calculate the similarity between the document d and the active cluster z as Among them, m z is the size of the document population in cluster z, D M is the number of active documents in the pattern, is the word frequency matrix of w in d, N d is the total number of words in d, N z is the total number of words in z, is the word frequency matrix of w in z, ICF w is the inverse clustering frequency, S w is the word specificity weight, V z∪d is V z ∪V d is the vocabulary of V z is the average vocabulary of z, V d is the average vocabulary of d, is the co-occurrence matrix of the document containing the frequency ratio between two adjacent terms, and δ is the adjacent window size.
5. A method for constructing a context-enhanced Dirichlet model that supports online clustering of short text streams, characterized in that: Highly specific terms have a smaller number of co-occurring neighboring data. Additionally, the normal neighboring data of a term depends on the defined window size. Therefore, the word specificity S of w i is defined as: where δ is the adjacent window size, and g(w i ) is defined as g(w i ) is used to calculate the ratio between the neighborhood population and the defined window size δ and its frequency in the document, where is the unique neighboring data of the term w i in the model, is the total term frequency, and the score is calculated by defining the boundary of the window size using the hyperbolic tangent sigmoid function.
Citation Information
Patent Citations
Malicious domain name cluster detection method and device based on word vector
CN113271292A
Regularized Latent Semantic Indexing for Topic Modeling
US20120330958A1