Short Text Clustering Method Based on Adaptive Variational Autoencoder

By introducing an adaptive variational autoencoder into short text clustering, high-dimensional text features are converted into low-dimensional separable features, the problem of poor clustering effect of short texts is solved and a more accurate clustering effect is achieved.

CN114625879BActive Publication Date: 2025-07-01BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210299111.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-13
Publication Date
2025-07-01
Estimated Expiration
2042-03-13

AI Technical Summary

Technical Problem

The high-dimensional features of short texts are difficult to convert into low-dimensional separable features, resulting in poor clustering effects, especially in high-dimensional data spaces, where traditional clustering methods have poor results.

Method used

The short text clustering method based on the adaptive variational autoencoder is adopted, and the text is converted into word vectors through Sentence-BERT, and pre-trained using the autoencoder, and iteratively optimized with K-means and the variational autoencoder to form a low-dimensional separable feature representation.

Benefits of technology

Effectively converting high-dimensional text features into low-dimensional separable features improves the accuracy and effectiveness of short text clustering, especially in high-dimensional data spaces, which significantly improves clustering results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114625879B_ABST
    Figure CN114625879B_ABST
Patent Text Reader

Abstract

The short text clustering method based on an adaptive variational autoencoder relates to the technical field of text clustering. First, the sentence-Bert method is used to represent short texts; secondly, an autoencoder is used to convert vectors into low-dimensional feature vectors, and the K-means method is used to extract clustering centers; then, the clustering centers are used as the expected means of the variational autoencoder to pre-train the input vectors, which are converted into feature vectors that satisfy the distribution with the clustering centers as the expected means; the feature vectors are used to construct a classifier according to the K-means algorithm, and the weights of the classifier and the encoder are fine-tuned through the classified distribution. Finally, the clustering results are obtained according to the fine-tuned encoder and classifier. The present invention can well handle the problem of high-dimensional sparsity of text vectors in short text clustering, and provides a new feature depth embedding algorithm for short text clustering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to text clustering technology, in particular to clustering of short texts and construction of corresponding deep algorithms. Background Art

[0002] With the rapid development of information technology, a large amount of short text data has been generated on various media platforms. Many fields such as news recommendation, user surveys, and event detection require mining valuable information from short text data. Compared with long texts, short texts have the characteristics of few words, ambiguity, and information non-standardization, which makes it difficult to extract and express text features.

[0003] Most words in short texts only appear once. Therefore, traditional vectorization methods based on word frequency cannot well express text features. Sparse feature representations bring problems such as insufficient word co-occurrence and lack of context information. To solve these problems, many word embedding models have been proposed, such as Word2vec, GloVe, ELMo, and BERT. These word embedding models are trained based on large corpora and use high-dimensional vectors to represent texts, which enriches the features of short texts to a certain extent and solves the problem of sparse text vectors, but also poses higher requirements for clustering algorithms.

[0004] Long-established clustering methods such as K-means and Gaussian Mixture Model (GMM) perform well in low-dimensional data spaces, but are less effective in high-dimensional data spaces. On the other hand, as an effective method for feature embedding, deep neural networks can map vectorized text data to a low-dimensional separable representation space, reducing the difficulty of text clustering algorithms.

[0005] Deep Embedded Clustering (DEC) combines clustering with deep embedding learning. It is a method proposed by Xie in 2016 for simultaneously learning feature representations and clustering assignments using a deep neural network. DEC maps data from the original space to a low-dimensional feature space, performs soft clustering assignment, and then iteratively optimizes the clustering objective, which also becomes the baseline algorithm for deep embedded clustering. However, when performing deep embedding, DEC generally uses an autoencoder (AE). The autoencoder uses the mean squared error (MSE) loss between the output and the input to optimize network parameters. It can obtain low-dimensional feature vectors containing original data information, but because the representation space is not regularized, it is easy to disrupt the data distribution, resulting in the phenomenon of different classes crossing and overlapping in the representation space. Summary of the Invention

[0006] To solve the above-mentioned problems, the present invention proposes a short text clustering method based on an adaptive variational autoencoder, aiming to convert high-dimensional features that can vectorize text into low-dimensional separable features, so as to accurately cluster short texts with similar semantics.

[0007] The short text clustering algorithm based on an adaptive variational autoencoder, and the steps of this method are as follows:

[0008] S1 Data collection;

[0009] S2 Input the text into sentence-Bert to convert it into word vectors;

[0010] S3 Use an autoencoder to pre-train the word vectors to obtain a dimensionality reduction encoder;

[0011] S4 Use K-means to cluster the data after dimensionality reduction to obtain the clustering labels and clustering centers of each text;

[0012] S5 Pre-train the text word vectors using a variational autoencoder, and use the clustering centers as the expected means to train the encoder network parameters;

[0013] S6 Use K-means to cluster the feature vectors generated by the pre-trained encoder to obtain the initial clustering centers;

[0014] S7 Soft-assign the vectors using the clustering centers;

[0015] S8 Use the auxiliary target distribution to learn and update the pre-trained encoder from the current high-confidence assignments and redefine the clustering centroids;

[0016] S9 Repeat S7 and S8, and when the convergence criterion or the number of iterations is met, output the clustering result. Description of the Drawings

[0017] Figure 1 It is a schematic diagram of the specific details of the short text clustering algorithm based on an adaptive variational autoencoder.

[0018] Figure 2 It is a flowchart of the short text clustering algorithm based on an adaptive variational autoencoder. Detailed Embodiment

[0019] The present invention proposes a short text clustering algorithm based on an adaptive variational autoencoder, and the main process of the method is as shown in the attached Figure 1 shown:

[0020] Combined with the attached Figure 1 The specific embodiments of the present invention will be described in detail:

[0021] Step S1, extract the text dataset. Extract the source text of microblogs from the microblog platform to construct a corpus of short texts D = {(s i , l i )|1 ≤ i ≤ n}, where n is the number of texts in the corpus D. S = {s1, s2,..., s n} represents the literal representation of all texts. L = {l1, l2,..., l n} indicates that the true label corresponds to s n . Here, since the unsupervised clustering method is adopted, the label is only used for the evaluation of the final result and does not participate in the model training process.

[0022] Step S2, without preprocessing the text data, use Sentence - BERT to represent the text in the vector space. Taking the i - th short text D i as an example, the text is represented as D i = {x i : x i ∈ R m}, where m is the dimension of the transformed sentence vector, and the generated sentence vector dimension is determined by the adopted model. Here, the model dimension is 384.

[0023] Step S3, use an auto - encoder to train the text vectors. For the transformed sentence vector x i ∈ R m . Construct an encoder to encode the original data:

[0024] Z i = f φ (x) = σ e (W e x i + b e ) ∈ R l #(1)

[0025] Then use a decoder to decode the original data:

[0026]

[0027] The loss function is to minimize the reconstruction error:

[0028]

[0029] where x i , and z i are the input data, output data, and latent variable respectively, f φ and g ψ represent the transformation functions of the encoder and decoder respectively. σ is the activation function, and here ReLU(x) is selected. W e and be are the weights and biases, where e and d represent the encoder and decoder respectively.

[0030] Autoencoders often update the network weights W e and biases b e by minimizing the reconstruction error. After completing the set number of iterations t, an encoder f φ (x): X ∈ R m → Z ∈ R l is obtained. t depends on the complexity of the network. In the present invention, t is set to 10, where Z is the latent feature space. Here, m is the dimension of the input sentence vector mentioned above, which is 384 dimensions, and l is the dimension of the hidden layer and is the same as the clustering target category k of the clustered text. Since the clustering category k is less than the input dimension d, a dimensionality reduction encoder f φ (x) is obtained.

[0031] Step S4, use K-means as the clustering algorithm to cluster the dimensionality-reduced text z i . Here, the Euclidean distance is used as the distance metric for the K-means algorithm. The goal of K-means is to select the centroid μ k in the cluster, which can minimize the within-cluster sum of squares:

[0032]

[0033] The purpose of this clustering step is to find the centroid and the text category k corresponding to each piece of text. In this paper, the expected mean of each text is obtained through the category k and the centroid and is denoted as

[0034] Step S5, adopt the network configuration of the deep autoencoder to deepen the network layers of the variational autoencoder. Use the expected mean μ * to train the variational autoencoder VAE. VAE aims to learn the generative model p(X, Z′) to maximize the marginal likelihood log p(X) of the dataset. Z′ is used to represent the representation space in VAE and is distinguished from the space in AE. Since it is difficult to calculate the integral of the latent variable, the marginal likelihood cannot be directly calculated. To solve this problem, VAE introduces a variational distribution q φ (Z′|X), and this variational distribution is approximated by a complex neural network parameter to approximate the true posterior distribution, and optimize the logp(X) of the evidence lower bound (ELBO):

[0035]

[0036] Where φ is the inference loss, θ is the decoder, the first term above is the reconstruction loss, and the second term is the KL divergence between the approximate posterior and the prior. In most VAEs, p(Z′) is a Gaussian distribution which is a common choice for the prior. The approximate posterior distribution q φ (Z′|X) and the prior p(Z′) can be calculated as:

[0037]

[0038] where μ i and σ i are the mean and variance of the approximate posterior distribution of the i-th dimensional representation space vector, respectively.

[0039] The present invention uses the clustering center μ in step 4 * as the expected mean of the feature distribution of the VAE, making p(Z′) become Therefore, the KL divergence can be calculated as:

[0040]

[0041] The second term of the VAE loss function is the difference from the ordinary autoencoder. The improvement based on the dataset here is mainly aimed at the KL divergence part. Without this term, it basically degenerates into a conventional AE, and the improvement of the present invention loses its effect, which is the phenomenon of KL divergence disappearance.

[0042] The present invention applies a fixed batch normalization (BN) to the output of the inference network μ i . This is a widely used regularization technique in deep learning. It can not only make the neuron output change normally but also be an effective method to prevent gradient explosion. Different from other tasks that apply BN to the hidden layer and seek fast and stable training, here BN is used as a tool to convert μ i into a distribution with fixed mean and variance. Mathematically, the regularized μ i is

[0043]

[0044] where and represent the approximate posterior before and after BN. μ Bi and σ Bi represent the mean and standard deviation of μ i , which are the biased estimates for each dimension of the sample. γ and β are the scale and shift parameters, and a fixed γ is used here.

[0045]

[0046] where τ ∈ (0, 1) is a constant, and in the present invention, τ = 0.5, while θ is a trainable parameter. Thus, the mean of the distribution of μ i is β and the variance is γ 2 . β is a learnable parameter that makes the distribution more flexible and is set to 0 in the present invention.

[0047] The improved variational autoencoder is called the self - adaptive variational autoencoder SVAE in the present invention. After pre - training with SVAE, the present invention takes the process from input to sampling as the encoder of SVAE. Through this encoder, the present invention performs a non - linear mapping f(X) on the data: X ∈ R d → Z′ ∈ R c , where the data dimension c is the same as the data dimension in the hidden layer of the autoencoder, and this setting refers to the number of clustering categories k.

[0048] Step S6: Use K - means to cluster the features Z′ in the feature space to obtain the cluster centers, and the purpose of this step is to serve as the initial weights of the DEC clustering layer.

[0049] Step S7: Use the cluster centers of K - means to calculate the soft cluster assignment of each data point of Z′ in the feature space. Use the single - degree - of - freedom t - distribution q ij to measure the similarity between the embedded point z i ′ and the centroid k j :

[0050]

[0051] where z i ′ = f(x i ) ∈ Z′ corresponds to the embedded x i ∈ X, where α is the degree of freedom of the student's t - distribution and α is taken as 1. And q ij can be interpreted as the probability of assigning sample i to cluster j (i.e., soft assignment).

[0052] Step S8: Use the auxiliary distribution p ij to improve the purity of clustering and to more strongly emphasize the data points with high credibility. The probability p ij in the auxiliary distribution P is calculated as follows:

[0053]

[0054] Fine - tune by matching the soft assignment and the target distribution. For this purpose, the present invention defines the objective as the KL divergence between the soft assignment and the auxiliary classification, as shown below:

[0055]

[0056] Step S9: Use Stochastic Gradient Descent (SGD) to jointly optimize the cluster centers k j For the parameters θ of the encoder in SVAE, the gradients of each sample and each cluster center are calculated as follows:

[0057]

[0058]

[0059] Use K-means as the weights for initializing the clustering layer, then use high-confidence predictions to determine the encoder and assign clusters. Repeat steps S7 and S8. When the number of iterations t1 reaches 2000, or when the class label change rate θ is less than 0.001, by taking the maximum value of the soft assignment q of each sample on the centroid, the assignment and clustering of the samples can be completed. Obtain the clustering result. The class label change rate θ is calculated as follows ij The maximum value can be used to complete the assignment and clustering of the samples. Obtain the clustering result. The class label change rate θ is calculated as follows

[0060]

[0061] L i And Are the labels of the i-th text and the previous label respectively, and n is the total number of samples.

[0062] Experimental Setup: The hardware environment used for the verification of this invention is but not limited to: The CPU is Inter Xeon4210R with a main frequency of 2.4 GHz, 64 GB of memory is used, two NVIDIA Gefore RTX 3060 graphics cards are used, and the operating system is windows 10. This invention uses the general sentence transformation library in Sentence-BERT to achieve the text vector representation of the general dataset (paraphrase-multilingual-MiniLM-L12-v2). The maximum sequence length of this model is 128, which can convert text into 384-dimensional vectors. The model size is 384 MB. It is a multilingual version based on the paraphrase-MiniLM-L12-v2 model and is trained on parallel data for more than 50 languages, mainly on multiple datasets such as AllNLI, sentence-compression, SimpleWiki, etc. During the pre-training process of this invention, Adam optimization (Kingma and Ba, 2015) is used, the batch_size is set to 64, the number of pre-training epochs is set to 15. The encoder on the SAE used during pre-training uses a network structure of [500, 500, 2000, 20], and the decoder is exactly the opposite, and there is also the same part in the VAE during the formal training. And during the training process in DEC, the SGD optimizer is used, a learning rate of 0.1 is used, a decay rate of 0.9 is set, the maximum number of iterations is 1500, and the batch size uses the same configuration as during pre-training.

[0063] The datasets used in the experiment include four English datasets for short text clustering and one Chinese dataset:

[0064] (1) SearchSnippets: A text collection of web search snippets, containing classifications of 8 different topics. (2) Stackoverflow: A collection of Q&A website posts, which was used as a dataset for a Kaggle challenge. This dataset contains question titles selected from 20 different categories. (3) Biomedical: A subset of the PubMed dataset, in which 20,000 paper titles were randomly selected from 20 groups. (4) Tweet: Consisting of 2,472 Tweet data with 89 categories. (5) In this paper, the ChineseNEWS Chinese public opinion dataset, which is a dataset used in actual projects, was used. Six typical events from 2017 to 2019 were searched, and the corresponding Weibo posts were crawled through event keywords. In this paper, 2,000 data were randomly selected from each event, obtaining a total of 12,000 datasets. To ensure the reliability of the dataset, the data was screened manually, and the 7th event "Weibo Night" and some texts unrelated to the event were marked and extracted as the 8th category. Therefore, a dataset of 8 categories was obtained in this paper.

[0065] The present invention is compared with the following clustering algorithms:

[0066] (1) Bow&TF-IDF: The sentences are vectorized by the frequency of words, and the sentences are converted into vectors with a dimension of 1,500, and the K-means algorithm is applied for clustering evaluation; (2) Sentence-BERT (SBERT): Through Sentence-BERT, the text is converted into vectors with a dimension of 384, which is also the text vectorization method used in this paper, and the K-means algorithm is applied for clustering evaluation; (3) VaDE: The product of combining the Gaussian mixture model and the variational autoencoder. It has a similar concept to the present invention, and both are first pre-trained through the autoencoder. The difference is that it obtains the initial prior of the data through GMM, and finally completes the feature embedding and clustering of the data through encoding, decoding, and backward updating of parameters; (4) STC2: Consists of three independent stages. For each dataset, it first pre-trains the words embedded in a large corpus using the word2Vec method. Then the convolutional neural network is optimized to further enrich the representation, and the sentences are put into K-means for the final stage of clustering; (5) Self-Train: Using SIF to enhance the pre-training word embeddings of Xu et al., following Xie et al. using a deep embedding clustering algorithm, which is an autoencoder obtained through hierarchical pre-training and then further optimized; (6) SCCL: The model consists of three components. For each dataset, the data includes the original data and the augmented data. After passing through a neural network to map the input data to the representation space, the contrast loss and the clustering loss are respectively applied to optimize the parameters of the encoder to complete the classification of the text.

[0067] The experimental results are shown in Table 1.

[0068] Table 1 Text Clustering Results

[0069]

[0070] The present invention shows the results of the algorithm in 5 datasets in Table 1. For the Chinese public opinion dataset used in the present invention, SVAE achieved the best results on this basis, leading by 3.1% in ACC compared to the excellent clustering algorithm Self-Train. This is because compared to AE, VAE can better improve the distribution of features in the latent space.

[0071] For the other 4 general datasets used in the present invention, it can be seen that SVAE has relatively good results in all three standard datasets. Among them, there is a huge improvement on StackOverflow. On the one hand, this is due to the improvement brought by the text vectorization model. It can be seen that SBERT also has a high accuracy on this dataset. On the other hand, SVAE improves the quality of feature embedding, bringing a higher improvement on this basis. Before this, it was difficult to obtain good clustering results for StackOverflow because of its large number of categories. The method mentioned in this paper brings a large improvement in accuracy because it makes good use of the cluster centers of the clustering algorithm.

[0072] Among them, SVAE mainly refers to the Self-Train algorithm, with an increase of 4.5% and 22.4% in ACC on SearchSnippets and StackOverflow respectively compared to Self-Train. The decrease in its Biomedical index is mainly because the general word training model used in this paper contains less content in this field, and the same conclusion can be drawn in SCCL. The method of SCCL leads by 3.6% in ACC on SearchSnippets. This is because in addition to the original dataset, SCCL also trains the model through an enhanced dataset. In addition to cluster optimization, it also uses contrastive loss to optimize the encoder, using a more complex architecture than SVAE. However, SVAE has an increase of 6.7% and 2% in ACC compared to SCCL on StackOverflow and Biomedical respectively. The results verify the effectiveness and importance of the framework proposed in this paper. In the context of the general language library, making full use of the prior clustering information improves the adaptability of the algorithm to different datasets.

Claims

1. Adaptive variational autoencoder-based short text clustering algorithm, characterized in that The steps are as follows: S1 Data collection; S2 Input the text into sentence-Bert to convert it into word vectors; S3 Use an autoencoder to pre-train the word vectors to obtain a dimensionality reduction encoder; S4 Use K-means to cluster the dimensionality-reduced data to obtain the cluster labels and cluster centers for each text; S5 Pre-train the text word vectors using a variational autoencoder, and use the cluster centers as the expected means to train the encoder network parameters; S6 Use K-means to cluster the feature vectors generated by the pre-trained encoder to obtain the initial cluster centers; S7 Use the cluster centers to perform soft assignment on the vectors; S8 Use the auxiliary target distribution to learn and update the pre-trained encoder and redefine the cluster centroids from the current high-confidence assignments; S9 Repeat S7 and S8, and when the convergence criterion or the number of iterations is met, output the clustering result; In step S2, there is no need to preprocess the data, and Sentence-BERT is used to represent the text in the vector space; In step S3, an autoencoder is used to train the text vector for the converted sentence vector x i ∈R m ; construct an encoder to encode the original data: z i = f φ (x) = σ e (W e x i + b e ) ∈ R l #(1) When using the decoder to decode the original data: The loss function is to minimize the reconstruction error: where x i , and z i are the input data, output data, and latent variable respectively, and f φ and g ψ represent the transformation functions of the encoder and decoder respectively; σ is the activation function, and ReLU(x) is selected here. W e and b e are the weights and biases, where e and d represent the encoder and decoder respectively; Autoencoders often update the network weights W by minimizing the reconstruction error e and the bias b e , and after completing the set number of iterations t, an encoder f φ (x): X ∈ R m → Z ∈ R l ; t is set to 10, where Z is the latent feature space, m here is the dimension of the input sentence vector mentioned above, which is 384 dimensions, l is the dimension of the hidden layer and is the same as the clustering target category k of the clustered text. Since the clustering category k is less than the input dimension d, a dimensionality reduction encoder f φ (x) is obtained; In step S4, the K-means is used as a clustering algorithm to cluster the text z after dimensionality reduction i ; here, the Euclidean distance is adopted as the distance metric of the K-means algorithm, and the goal of K-means is to select the centroid μ in the cluster k , which can minimize the within-cluster sum of squares: The purpose of this step of clustering is to find the centroid and the text category k corresponding to each piece of text; through the category k and the centroid the expected mean value of each text is obtained and denoted as After preprocessing, the vectorized X of the text and the clustering center μ of the dimensionality-reduced text are obtained. * ; In step S7, calculate the soft cluster assignment of each data point Z' in the feature space according to the cluster center; use a single-degree-of-freedom t-distribution q ij , to measure the embedded point z i ′ and the centroid k j : between the similarities where z i ′ = f(x i ) ∈ Z′ corresponds to x after SVAE embedding i ∈ X, where α is the degree of freedom of the student's t - distribution, and q ij is the probability of assigning sample i to cluster j, i.e., soft assignment, and α takes 1.

2. The short text clustering algorithm based on the adaptive variational encoder according to claim 1, wherein: Step S5, use the expected mean μ * Train the variational autoencoder VAE, and add a BN layer to the variational autoencoder VAE to prevent the KL divergence in the VAE loss function from disappearing, jointly constituting the framework of the SVAE; Use the cluster center μ in step S4 * As the expected mean of the feature distribution of the VAE, make p(Z′) become Therefore, the KL divergence is calculated as: Using BN as a tool to transform μ i into a distribution with fixed mean and variance; mathematically, the regularized μ i is Among them and represent the approximate posterior before and after BN; μ Bi and σ Bi represent the mean and standard deviation of μ i which are the biased estimates for each dimension of the sample, and γ and β are the shift parameters of the scale where τ ∈ (0, 1) is a constant and θ is a trainable parameter; thus, the mean of the distribution of μ i is β and the variance is γ 2 ; β is a learnable parameter.

3. The short text clustering algorithm based on the adaptive variational encoder according to claim 1, wherein: In step S6, use K-means to cluster the features Z′ in the SVAE representation space to obtain the cluster centers, and the purpose of this step is to be used as the initial weights of the DEC clustering layer; In step S7, calculate the soft cluster assignment of each data point Z' in the feature space according to the cluster center; use a single-degree-of-freedom t-distribution q ij , to measure the embedded point z i ' and the centroid k j : between similarities where z i ′ = f(x i ) ∈ Z′ corresponds to x after SVAE embedding i ∈ X, where α is the degree of freedom of the student's t-distribution, and q ij is the probability of assigning sample i to cluster j, i.e., soft assignment, and α is taken as 1; Step S8: Use the auxiliary distribution p ij , to improve the purity of clustering and to place more emphasis on data points with high confidence; the probability p in the auxiliary distribution P ij is calculated as follows: Fine-tune through soft assignment and target distribution matching. For this purpose, the objective is defined as the KL divergence between the soft assignment and the auxiliary classification, as shown below: Step S9: Use stochastic gradient descent to jointly optimize the cluster centers k cntj For the parameters θ of the encoder in SVAE, the gradients for each sample and each cluster center are calculated as follows: Use K-means as the weights of the initialization clustering layer, and then use high-confidence predictions to determine the encoder and assign clusters. Repeat steps S7 and S8. When the iteration count t1 reaches 2000, or when the class label change rate θ is less than 0.001, the samples are assigned and clustered by taking the maximum value of the soft assignment q of each sample on the centroid ij to obtain the clustering result; the class label change rate θ is calculated as follows L i and are respectively the label of the i-th text and the previous label, and n is the total number of samples.

Citation Information

Patent Citations

  • Short text topic recognition method based on Dirichlet variational auto-encoder

    CN112597769A

  • Multi-modal adaptive fusion depth clustering model and method based on auto-encoder

    CN112884010A