Short text clustering method for exploring higher semantic quality for low-confidence data

Through mask language modeling and pseudo-label algorithm combined with comparison learning, the problems of low feature representation quality and high cluster crossover in short text clustering are solved, and the short text clustering effect with higher semantic quality is achieved.

CN120541232APending Publication Date: 2025-08-26UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510626605.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the prior art, the problem of text characteristics indicating low data purity and excessive crossing between clusters in short text clusters is difficult to effectively separate short text clusters, and the lack of stable negative examples in comparison learning leads to poor results.

Method used

Sentence-BERT is pre-trained using mask language modeling, combining pseudo-label algorithm and semi-supervised learning, and by comparative learning, the distance between outstrips and non-outside points is pulled closer by using cluster heads to improve feature representation quality.

Benefits of technology

The semantic quality of short text clustering is improved, smaller in-cluster distances and larger intercluster distances are generated, and the generalization performance and clustering effect of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541232A_ABST
    Figure CN120541232A_ABST
Patent Text Reader

Abstract

The invention provides a short text clustering semantic word vector generation method for exploring higher semantic quality for low-confidence data, and belongs to the technical field of deep clustering. The method comprises the following steps: pre-training a Sensor-BERT, and obtaining a pre-trained Sensor-BERT which is used as a feature extractor of a data set; extracting feature vectors of the short text, the weak enhancement text and the strong enhancement text of the data set by using a feature extractor; calculating pseudo labels and detecting outliers, and dividing the data set into outlier data and non-outlier data; calculating cluster head loss, calculating contrast head loss, and calculating total training loss; continuously updating the parameters of the pre-trained Sensor-BERT, the clustering head and the contrast head, and obtaining the trained Sensor-BERT and the clustering head; and calculating a clustering result by using the trained clustering head. According to the method, the data relationship of the text data set and the difference of the cluster level are comprehensively considered, and the semantic quality is mined from the text as much as possible, so that the generated text feature representation has better separability, and a very good clustering effect is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep clustering, and in particular relates to a short text clustering method for exploring higher semantic quality for low-confidence data. Background Art

[0002] Short text clustering (STC) is a key task in unsupervised learning. Its goal is to cluster unlabeled short texts. With the development of technology and the popularity of social media, short texts such as online comments, Weibo posts, and search terms are growing rapidly. Therefore, organizing these texts according to specific topics or event discussions is a critical and important step in data mining. Tasks such as text summarization, public opinion analysis, and event monitoring all require clustering.

[0003] Because short texts have relatively little semantic information, high noise, and high dimensionality, traditional clustering methods perform poorly on short text clustering tasks. Traditional clustering methods often rely on distance matrices in the data space. However, obtaining a good distance matrix often requires extracting better feature vectors from short texts. To address this issue, previous deep learning methods used neural networks to enrich sparse feature representations to obtain better feature vectors. At the same time, some researchers have also attempted to add additional knowledge to short texts to enrich text representations and obtain better feature vectors.

[0004] However, even with the use of neural networks, the generated text feature representation still suffers from low data purity and high overlap between clusters. To address the low quality of the initially extracted text feature vectors, the present invention uses Mask Language Modeling (MLM) to pre-train Sentence-BERT on a specific dataset to improve Sentence-BERT's initial text feature representation and increase its purity. However, the overlap between clusters is still significant.

[0005] On the other hand, contrastive learning—a paradigm that teaches models to compare—has been highly successful in the field of self-supervised learning, achieving groundbreaking success in image-text clustering and sentence representation tasks. Contrastive learning can disperse data feature representations across the data space to alleviate the problem of overlap between data clusters. Its basic idea is to bring the feature vectors of positive pairs of text data closer together while pulling the feature vectors of negative pairs farther apart. Positive pairs are augmented from the same source text, while negative pairs are augmented from different source texts. However, due to the lack of true data labels, inherent false negative pairs in contrastive learning can make contrastive learning unstable or even yield poor results. These inherent false negative pairs occur when the source text of the two augmented texts in a negative pair belongs to the same semantic cluster. Therefore, it is impossible to pull the feature vectors of such texts farther apart because they belong to the same semantic cluster and should be closer.

[0006] To solve this problem, the present invention introduces a pseudo-label mechanism, which is combined with contrastive learning to provide certain supervisory information guidance for feature learning, helping to generate higher quality negative sample pairs for self-supervised learning. In order to ensure the quality of the pseudo-labels, it is necessary to detect outliers in each cluster and discard their pseudo-labels. However, simply ignoring all outlier data will damage the learning process. In order to further explore the semantic quality of the text, the present invention combines the algorithm with semi-supervised learning (SSL). Semi-supervised learning focuses on learning with a small amount of labeled data and a large amount of unlabeled data. It has shown great potential in practical applications because it greatly reduces the need for time-consuming annotation. Its main challenge lies in how to effectively utilize the information of unlabeled data to improve the generalization performance of the model. The SoftMatch method derives a truncated Gaussian function to fit the confidence distribution by assuming the distribution of pseudo-labels. According to the deviation of its confidence from the Gaussian mean, a lower weight is assigned to pseudo-labels that may have low confidence, so as to maintain a high number and high quality of pseudo-labels during the training process. Inspired by this idea, the present invention uses a similar sample weight function to weight outlier data and mines the semantic quality of the text as much as possible. Summary of the Invention

[0007] The present invention aims to provide a short text clustering method for exploring higher semantic quality for low-confidence data. The method first uses masked language modeling to improve the quality of Sentence-BERT text feature vectors for special data sets, then uses a pseudo-label algorithm to improve the quality of negative pairs in contrastive learning, uses contrastive learning to shorten the distance between outliers and non-outliers, and uses the idea of ​​semi-supervised learning to improve the mining of semantic quality in outlier data during clustering, and shortens the distance between non-outliers through clustering heads. In this way, smaller intra-cluster distances and larger inter-cluster distances are obtained. This solves the technical problems in the prior art of low purity of text feature representation data, high crossover between clusters, and poor clustering effect.

[0008] In order to solve the above technical problems, the specific technical solutions of the present invention are as follows:

[0009] A method for generating short text clustering semantic word vectors for exploring higher semantic quality for low-confidence data, the method comprising the following steps:

[0010] Step S1: Use masked language modeling to pre-train Sentence-BERT to obtain a pre-trained Sentence-BERT, which is used as a feature extractor for the dataset;

[0011] Step S2: Use the pre-trained Sentence-BERT as a feature extractor to extract feature vectors of the short text, weakly enhanced text, and strongly enhanced text in the dataset;

[0012] Step S3: Build a clustering head, calculate pseudo labels and detect outliers, divide the dataset into outlier data and non-outlier data; train Sentence-BERT, and update the parameters and pseudo labels of the pre-trained Sentence-BERT;

[0013] Step S4: Calculate the outlier loss and non-outlier loss, and then calculate the clustering head loss; construct the comparison head, calculate the weak enhancement text loss and strong enhancement text loss, and then calculate the comparison head loss; calculate the total training loss; based on the pseudo-labels, continuously update the parameters of the pre-trained Sentence-BERT, clustering head, and comparison head through backpropagation and stochastic gradient descent to obtain the trained Sentence-BERT, clustering head, and comparison head;

[0014] Step S5: Calculate the clustering results using the trained clustering head.

[0015] Furthermore, in step S2, the data set contains N short texts, represented as where x i Represents the i-th short text;

[0016] For the dataset Short text x in i , based on the context, use the pre-trained T5 model to predict 20% of the random mask segments, and then select the top 5 segments with the largest probability to replace them to get the short text x i Weakly enhanced text Dataset The weakly enhanced text corresponding to each short text in the composes a weakly enhanced text dataset

[0017] For the dataset Short text x in i , replace 20% of the words with synonyms, delete 20% of the words randomly, and exchange 20% of the words randomly to get the short text x i Strongly enhanced text Dataset The strongly enhanced text corresponding to each short text in the text constitutes a strongly enhanced text dataset

[0018] Use pre-trained Sentence-BERT as a feature extractor to extract feature vectors for short text, weakly enhanced text, and strongly enhanced text;

[0019] Short text x i The eigenvector of i =f′ θ (x i ); weakly enhanced text The eigenvector of Strongly enhanced text The eigenvector of f′ θ (·) denotes pre-trained Sentence-BERT.

[0020] Furthermore, in step S3, pseudo label generation is transformed into an optimal transmission problem to be solved, and the pseudo label set is obtained by solving the optimal transmission problem using the Sinkhorn algorithm. where y′ i ∈{0,1,…,K-1}, represents a short text x i The corresponding pseudo label is the category number assigned to the short text; then the isolation forest outlier detection algorithm is used to detect the outlier data in each class, and the pseudo label of the outlier data is set to -1, thereby Divided into two groups: non-outlier data and outlier data

[0021] Furthermore, calculating the outlier loss and the non-outlier loss in step S4 and then calculating the clustering head loss includes the following steps:

[0022] For each iteration, from the dataset A batch of size M is randomly selected from and separated into non-outlier batches and outlier batches

[0023] For non-outlier batches Based on the pseudo labels obtained previously, a cross entropy loss function is used to calculate the non-outlier loss. Calculated as follows:

[0024]

[0025] in, represents the batch size of non-outliers, Represents short text x i The category probability distribution vector, p i [k] represents the short text x i The probability of belonging to the k+1th category, y′ i Represents short text x i The corresponding pseudo label, p i [y′ i ] represents the short text x i Belongs to the y′th i +1 probability of the class.

[0026] For outlier batches A weighted cross entropy function is used to calculate the outlier loss, the outlier loss Calculated as follows:

[0027]

[0028] in, represents the outlier batch size, Represents short text x i Corresponding weakly enhanced text The category probability distribution vector, Represents short text x i Corresponding strongly enhanced text The category probability distribution vector of Indicates strongly enhanced text The probability of belonging to the k+1th category, Represents the clustering head for weakly enhanced text The predicted label, argmax kIndicates the index corresponding to the maximum value, Indicates strongly enhanced text Belongs to the y′th i +1 category probability, β(·) is a sample weight function.

[0029] The overall loss of the clustering head is the sum of non-outlier loss and outlier loss, clustering head loss Calculated as follows:

[0030]

[0031] Furthermore, in step S4, a comparison head is constructed, and the weakly enhanced text loss and the strongly enhanced text loss are calculated, and then the comparison head loss is calculated. The calculation of the total training loss includes the following steps:

[0032] The comparison head is a 3-layer perceptron with 768, 768, and 128 nodes in each layer, respectively. ReLU is used as the activation function between layers.

[0033] For each iteration, from the dataset Randomly select a batch of size M from the short text batch And from weakly enhanced text data and strongly enhanced text data Select the corresponding data and record them as weakly enhanced text batches and strongly enhanced text batches For weakly augmented text batches Weakly enhanced text in Calculating weakly enhanced text loss The weakly enhanced text loss is calculated as follows:

[0034]

[0035] in, Represents and short text x i The index set of all data that constitute the negative pairs, and the false negative pairs are discarded according to the pseudo labels, y′ i Represents short text x i The pseudo label, y′ j Represents other short text x j The pseudo labels of is the cosine similarity, ‖·‖2 represents the L2 norm, and τ is the temperature parameter;

[0036] For strongly enhanced text batches Strongly enhanced text in Computing the Strong Enhanced Text Loss The strong enhanced text loss is calculated as follows:

[0037]

[0038] The overall loss of the comparison head It is the average of the strong enhancement text loss of the strong enhancement text dataset and the weak enhancement text loss of the weak enhancement text dataset. The contrast head loss is calculated as follows:

[0039]

[0040] Calculate the total training loss The calculation expression is:

[0041]

[0042] Where λ(l) represents a monotonically decreasing function.

[0043] Compared with the prior art, the present invention has the following beneficial technical effects:

[0044] 1) The present invention more fully mines the semantic quality of low-confidence data and improves the generalization performance of the model.

[0045] 2) The present invention adopts a more appropriate weak enhancement scheme, which improves the quality of weakly enhanced text and further improves the model performance.

[0046] 3) The present invention comprehensively considers the data relationship at the instance level and the distinctiveness at the cluster level of the text dataset, and mines the semantic quality from the text as much as possible, so that the generated text feature representation has better separability and achieves a good clustering effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0048] Figure 1 Schematic diagram of the short text clustering model framework for exploring higher semantic quality for low-confidence data according to the present invention.

[0049] Figure 2 This is the pseudo code of the short text clustering algorithm of the present invention for exploring higher semantic quality for low-confidence data. DETAILED DESCRIPTION

[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0051] This paper proposes a ShortText Clustering Exploring Higher Semantic Quality Toward Low-Confidence Data (TC-SQL), such as Figure 1 As shown, the method includes the following steps:

[0052] Step S1: Use masked language modeling to pre-train Sentence-BERT (Sentence Bidirectional Encoder Representations from Transformers) to obtain pre-trained Sentence-BERT. The pre-trained Sentence-BERT is used as the feature extractor of the dataset.

[0053] For a given dataset containing N short texts where x i Represents the i-th short text, and uses masked language modeling to pre-train Sentence-BERT to obtain the dataset feature extractor.

[0054] like Figure 1 As shown in the figure, in the pre-training phase, 15% of the words in the short text in the dataset are randomly masked and the corresponding word list id of the words is recorded. Then, Sentence-BERT is used to extract the feature vector of the text, and a classifier is used to predict the masked words based on the feature vector. Using an unsupervised method to pre-train Sentence-BERT, a Sentence-BERT can obtain a good initial feature vector. The version of Sentence-BERT used is distilbert-base-nli-stsb-mean-token. The pre-trained Sentence-BERT will be used in subsequent training.

[0055] Step 2: Use the pre-trained Sentence-BERT as a feature extractor to extract feature vectors of short text, weakly enhanced text, and strongly enhanced text in the dataset.

[0056] For the dataset Short text x in i , based on the context, use the pre-trained T5 model (Text to Text Transfer Transformer) to predict 20% of the random mask segments, and then select the top 5 probability segments for replacement to obtain the short text x i Weakly enhanced text Dataset The weakly enhanced text corresponding to each short text in the composes a weakly enhanced text dataset Traditional BERT-based contextual text augmentation works at the word level, masking the words in a sentence and then predicting the fill-in. The T5 model, however, masks text segments rather than individual words. This preserves the meaning of phrases and allows the model to better understand them. Furthermore, masking focuses on including the entity portion of a sentence or the portion describing time, making this data augmentation method more suitable for social media text.

[0057] For the dataset Short text x in i , replace 20% of the words with synonyms, delete 20% of the words randomly, and exchange 20% of the words randomly to get the short text x i Strongly enhanced text Dataset The strongly enhanced text corresponding to each short text in the text constitutes a strongly enhanced text dataset

[0058] Note that Sentence-BERT before and after pre-training is f θ (·) and f′ θ (·) We use the pre-trained Sentence-BERT as a feature extractor to extract feature vectors for short text, weakly enhanced text, and strongly enhanced text.

[0059] Short text x i The eigenvector of i =f′ θ (x i ); weakly enhanced text The eigenvector of Strongly enhanced text The eigenvector of in

[0060] Step S3: Build a clustering head, calculate pseudo labels and detect outliers, divide the dataset into outlier data and non-outlier data; and train Sentence-BERT, updating the parameters and pseudo labels of the pre-trained Sentence-BERT.

[0061] Clustering head g C (·): The clustering head is a three-layer perceptron. The number of nodes in each layer is 768, 768, and K, respectively. ReLU is used as the activation function between layers, where K is the number of clusters into which the dataset is to be clustered.

[0062] Use pre-trained Sentence-BERT as a feature extractor for the dataset Perform feature extraction to obtain a data set The eigenvector of Pseudo-label generation can be transformed into an optimal transport problem to solve, that is, to calculate the transfer matrix from sample distribution to category distribution. The sample distribution is denoted as Category distribution is recorded as The transfer matrix is ​​denoted as The feature vector H is mapped by the clustering head to obtain the probability matrix Expressed as P = g C (H), where g C (·) represents the clustering head. Taking the negative logarithm of the probability matrix P, we can get the input cost matrix of the optimal transmission. Expressed as C = -logP. The optimal transmission problem to be solved is expressed as follows:

[0063]

[0064] stT1=a,T T 1=b,T≥0,b T 1=1

[0065] Among them, T ij represents the element in the i-th row and j-th column of the transmission matrix T, C ij Represents the element in the i-th row and j-th column of the input cost matrix C, ∈1 is a regularization hyperparameter, H(T) = ∑ i,j (T ij log T ij -T ij ) is entropy regularization, ∈2 represents a penalty term balancing hyperparameter, 1 represents a full 1 vector, is the sample distribution, b is the unknown category distribution, and Ψ(b) = -logb-log(1-b) is the penalty term for b.

[0066] Use the Sinkhorn algorithm to solve the above equation to obtain the transfer matrix Then calculate the pseudo label set where y′ i =argmax j T i ∈{0,1,…,K-1},T i Table represents the i-th row of the transmission matrix T, y′ i Represents short text x i The corresponding pseudo label, that is, the category number assigned, argmax j Represents the index corresponding to the maximum value. Then the Isolation Forest outlier detection algorithm is used to detect the outlier data in each class and the pseudo label of the outlier data is set to -1, thereby Divided into two groups: non-outlier data and outlier data During training, the parameters and pseudo-labels of Sentence-BERT are updated. The pseudo-labels are updated every 50 training batches. and outlier data It will also be updated synchronously.

[0067] Step S4: Calculate the outlier loss and non-outlier loss, and then calculate the clustering head loss; construct the comparison head, calculate the weak enhancement text loss and strong enhancement text loss, and then calculate the comparison head loss; calculate the total training loss; based on the pseudo-label, continuously update the parameters of the pre-trained Sentence-BERT, clustering head, and comparison head through backpropagation and stochastic gradient descent to obtain the trained Sentence-BERT and clustering head.

[0068] The following describes the calculation of the clustering head loss value and the network structure and loss value calculation of the comparison head:

[0069] The clustering head is also used to classify the short text x i The eigenvector h i Mapped to a K-dimensional subspace, represented as p i Represents short text x i The clustering head is also used to weakly enhance the text The eigenvector of Mapped to a K-dimensional subspace, represented as Indicates weakly enhanced text The clustering head is also used to strongly enhance the text The eigenvector of Mapped to a K-dimensional subspace, represented as Indicates strongly enhanced text The class probability distribution vector of .

[0070] For each iteration, from the dataset A batch of size M is randomly selected from and separated into non-outlier batches and outlier batches

[0071] For non-outlier batches Based on the pseudo labels obtained previously, a cross entropy loss function is used to calculate the non-outlier loss. Calculated as follows:

[0072]

[0073] in, represents the batch size of non-outliers, Represents short text x i The category probability distribution vector, p i [k] represents the short text x i The probability of belonging to the k+1th category, y′ i Represents short text x i The corresponding pseudo label, p i [y′ i ] represents the short text x i Belongs to the y′th i +1 probability of the class.

[0074] Outlier data are considered as difficult learning samples with low confidence. Previous work either used all samples indiscriminately or discarded low-confidence samples in learning, thus encountering performance bottlenecks. This paper further mines semantic quality from these data to improve the generalization performance of the model. A weighted cross entropy function is used to calculate the outlier loss, the outlier loss Calculated as follows:

[0075]

[0076] in, represents the outlier batch size, Represents short text x i Corresponding weakly enhanced text The category probability distribution vector of Represents short text x i Corresponding strongly enhanced text The category probability distribution vector, Indicates strongly enhanced text The probability of belonging to the k+1th category, Represents the clustering head for weakly enhanced text The predicted label, argmax k Indicates the index corresponding to the maximum value, Indicates strongly enhanced text Belongs to the y′th i +1 category probability, β(·) is a sample weight function with a range of [0,β max ], defined as follows:

[0077]

[0078] The sample weight function is a dynamic truncated Gaussian distribution, where p is the input probability vector, μ t and are the mean and variance of the Gaussian distribution at the tth iteration, β max UA(·) is a hyperparameter that controls the maximum value of the sample weight. It is a “uniform alignment” method used to make data of different categories have more uniform pseudo labels.

[0079] To estimate μ t and , computing the predicted statistics at each iteration:

[0080]

[0081] in, represents the Gaussian mean estimator of the current outlier batch, represents the sample mean of the current outlier batch, represents the Gaussian variance estimator of the current outlier batch, represents the sample variance of the current outlier batch, Indicates the outlier batch size.

[0082] Then use the Exponential Moving Average (EMA) and historical data, and introduce the momentum parameter m to update the mean and variance of the outliers:

[0083]

[0084] in, and are the estimated values ​​of the mean and variance of the Gaussian distribution at the tth iteration, and are the estimated values ​​of the Gaussian mean and variance at the t-1th iteration, Indicates the outlier batch size. Here, unbiased variance is used for EMA, and the initial mean and variance are set to and m is the momentum parameter.

[0085] The expression for "even alignment" is:

[0086]

[0087] Wherein, Normalize(·)=(·) / ∑(·) represents a regularization function, which is used to ensure that the sum of the confidence distribution is 1. is a uniformly distributed probability vector, Represents the sample mean of the current outlier batch.

[0088] The overall loss of the clustering head is the sum of non-outlier loss and outlier loss, clustering head loss Calculated as follows:

[0089]

[0090] Contrast headg I (·): The comparison head is a 3-layer perceptron with 768, 768, and 128 nodes in each layer, respectively. ReLU is used as the activation function between layers. The function of the comparison head is to map the feature vectors of weakly enhanced text and strongly enhanced text into a 128-dimensional subspace, which can be formally expressed as and in Represents the weakly enhanced sample of the i-th short text The eigenvector of is the vector after mapping, Represents the strong enhancement sample of the i-th short text The eigenvector of is the vector after mapping.

[0091] will be represented by the same short text x i Weakly enhanced text and enhanced text Constitute a positive example pair By short text x i Weakly enhanced text and another short text x j Enhanced strong enhanced text Negative pairs Where i≠j.

[0092] The function of the contrast head is to bring the feature vectors of the positive pairs closer together and the feature vectors of the negative pairs farther apart, so that Sentence-BERT can obtain more discriminative feature vectors. However, due to the lack of true labels for these data, the effect of contrastive learning is affected by inherent false negative pairs. The two different short text data of the inherent false negative pair come from the same semantic cluster class, so the feature vectors of the inherent false negative pair should not be pulled apart. Therefore, the present invention uses the pseudo labels generated in step S3 to alleviate the impact of inherent false negative pairs.

[0093] For each iteration, from the dataset Randomly select a batch of size M from the short text batch And from weakly enhanced text data and strongly enhanced text data Select the corresponding data and record them as weakly enhanced text batches and strongly enhanced text batches

[0094] For weakly augmented text batches Weakly enhanced text in Calculating weakly enhanced text loss The weakly enhanced text loss is calculated as follows:

[0095]

[0096] in, Represents and short text x i The index set of all data that constitute the negative pairs, and the false negative pairs are discarded according to the pseudo labels, y′ i Represents short text x i The pseudo label, y′ j Represents other short text x j The pseudo labels of is the cosine similarity, ‖·‖2 represents the L2 norm, and τ is the temperature parameter.

[0097] For strongly enhanced text batches Strongly enhanced text in Computing the Strong Enhanced Text Loss The strong enhanced text loss is calculated as follows:

[0098]

[0099] The overall loss of the comparison head It is the average of the strong enhancement text loss of the strong enhancement text dataset and the weak enhancement text loss of the weak enhancement text dataset. The contrast head loss is calculated as follows:

[0100]

[0101] The present invention trains the clustering head and the comparison head together and uses a function to dynamically adjust The weights in the joint training. The total training loss The calculation expression is:

[0102]

[0103] Among them, λ(l) represents a monotonically decreasing function with a range of [5,15], which can be expressed as:

[0104]

[0105] Among them, L represents the total number of iteration steps, l represents the current number of iteration steps, l∈[0,L].

[0106] Based on the pseudo-labels, the parameters of the pre-trained Sentence-BERT, clustering head, and comparison head are continuously updated through backpropagation and stochastic gradient descent to obtain the trained Sentence-BERT, clustering head, and comparison head.

[0107] Step S5: Calculate the clustering results using the trained clustering head.

[0108] After training, clustering results are calculated using clustering heads. i First, use the pre-trained Sentence-BERT to obtain the feature vector h of the short text i , and then use the clustering head g C (·) Obtain the category probability distribution vector of the short text, and finally select the index corresponding to the maximum probability as the short text x i The cluster label c i .

[0109] The entire training process uses Figure 2 The pseudo code shown is summarized in the following figure. The value of the temperature parameter τ in step S4 is set to 1.0. The learning rate for pre-training Sentence-BERT is set to 5×10 -6 , the learning rate of the contrast head and clustering head is set to 5×10 -5 , the batch size M = 400, and the total number of iterations L = 1000. The present invention uses a masked pre-trained language model, contrastive learning based on a non-outlier pseudo-labeling algorithm, and clustering learning based on semi-supervised learning to simultaneously consider the relationship between data in the dataset and the distinction between cluster semantics, generating high-quality text feature vectors and improving the short text clustering effect.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A short text clustering semantic word vector generation method for exploring higher semantic quality for low-confidence data, characterized by: The method comprises the following steps: Step S1: Use masked language modeling to pre-train Sentence-BERT to obtain a pre-trained Sentence-BERT, which is used as a feature extractor for the dataset; Step S2: Use the pre-trained Sentence-BERT as a feature extractor to extract feature vectors of the short text, weakly enhanced text, and strongly enhanced text in the dataset; Step S3: Build a clustering head, calculate pseudo labels and detect outliers, divide the dataset into outlier data and non-outlier data; train Sentence-BERT, and update the parameters and pseudo labels of the pre-trained Sentence-BERT; Step S4: Calculate the outlier loss and non-outlier loss, and then calculate the clustering head loss; construct the contrast head, calculate the weak enhancement text loss and strong enhancement text loss, and then calculate the contrast head loss; Calculate the total training loss; based on the pseudo-labels, continuously update the parameters of the pre-trained Sentence-BERT, clustering head, and comparison head through backpropagation and stochastic gradient descent to obtain the trained Sentence-BERT, clustering head, and comparison head; Step S5: Calculate the clustering results using the trained clustering head.

2. The method for generating short text clustering semantic word vectors for exploring higher semantic quality for low-confidence data according to claim 1, characterized in that: In step S2, The dataset contains N short texts, represented as where x i Represents the i-th short text; For the dataset Short text x in i , based on the context, use the pre-trained T5 model to predict 20% of the random mask segments, and then select the top 5 segments with the largest probability to replace them to get the short text x i Weakly enhanced text Dataset The weakly enhanced text corresponding to each short text in the composes a weakly enhanced text dataset For the dataset Short text x in i , replace 20% of the words with synonyms, delete 20% of the words randomly, and exchange 20% of the words randomly to get the short text x i Strongly enhanced text Dataset The strongly enhanced text corresponding to each short text in the text constitutes a strongly enhanced text dataset Use pre-trained Sentence-BERT as a feature extractor to extract feature vectors for short text, weakly enhanced text, and strongly enhanced text; Short text x i The eigenvector of i =f′ θ (x i ); weakly enhanced text The eigenvector of Strongly enhanced text The eigenvector of f′ θ (·) denotes pre-trained Sentence-BERT.

3. The method for generating short text clustering semantic word vectors for exploring higher semantic quality for low-confidence data according to claim 2, characterized in that: In step S3, The pseudo-label generation is transformed into an optimal transmission problem for solution. The pseudo-label set is obtained by solving the optimal transmission problem using the Sinkhorn algorithm. where y′ i ∈{0,1,…,K-1}, represents a short text x i The corresponding pseudo label is the category number assigned to the short text; then the isolation forest outlier detection algorithm is used to detect the outlier data in each class, and the pseudo label of the outlier data is set to -1, thereby Divided into two groups: non-outlier data and outlier data 4. The method for generating short text clustering semantic word vectors for exploring higher semantic quality for low-confidence data according to claim 3, characterized in that: Calculating the outlier loss and non-outlier loss in step S4, and then calculating the clustering head loss, includes the following steps: For each iteration, from the dataset A batch of size M is randomly selected from and separated into non-outlier batches and outlier batches For non-outlier batches Based on the pseudo labels obtained previously, a cross entropy loss function is used to calculate the non-outlier loss. Calculated as follows: in, represents the batch size of non-outliers, Represents short text x i The category probability distribution vector, p i [k] represents the short text x i The probability of belonging to the k+1th category, y′ i Represents short text x i The corresponding pseudo label, p i [y′ i ] represents the short text x i Belongs to the y′th i +1 probability of the class; For outlier batches A weighted cross entropy function is used to calculate the outlier loss, the outlier loss Calculated as follows: in, represents the outlier batch size, Represents short text x i Corresponding weakly enhanced text The category probability distribution vector of Represents short text x i Corresponding strongly enhanced text The category probability distribution vector of Indicates strongly enhanced text The probability of belonging to the k+1th category, Represents clustering head for weakly enhanced text The predicted label, argmax k Indicates the index corresponding to the maximum value, Indicates strongly enhanced text Belongs to the y′th i +1 category probability, β(·) is a sample weight function; The overall loss of the clustering head is the sum of non-outlier loss and outlier loss, clustering head loss Calculated as follows:

5. The method for generating short text clustering semantic word vectors for exploring higher semantic quality for low-confidence data according to claim 3, characterized in that: In step S4, a comparison head is constructed, and the weakly enhanced text loss and the strongly enhanced text loss are calculated. Then, the comparison head loss is calculated. The calculation of the total training loss includes the following steps: The comparison head is a 3-layer perceptron with 768, 768, and 128 nodes in each layer, respectively. ReLU is used as the activation function between layers. For each iteration, from the dataset Randomly select a batch of size M from the short text batch And from weakly enhanced text data and strongly enhanced text data Select the corresponding data and record them as weakly enhanced text batches and strongly enhanced text batches For weakly augmented text batches Weakly enhanced text in Calculating weakly enhanced text loss The weakly enhanced text loss is calculated as follows: in, Represents and short text x i The index set of all data that constitute the negative pairs, and the false negative pairs are discarded according to the pseudo labels, y′ i Represents short text x i The pseudo label, y′ j Represents other short text x j Pseudo labels, is the cosine similarity, ‖·‖2 represents the L2 norm, and τ is the temperature parameter; For strongly enhanced text batches Strongly enhanced text in Computing the Strong Enhanced Text Loss The strong enhanced text loss is calculated as follows: The overall loss of the comparison head It is the average of the strong enhancement text loss of the strong enhancement text dataset and the weak enhancement text loss of the weak enhancement text dataset. The contrast head loss is calculated as follows: Calculate the total training loss The calculation expression is: Where λ(l) represents a monotonically decreasing function.