Short text clustering method, device and equipment based on large language model assistance and storage medium
By introducing large language models and information entropy sampling technology into short text clustering, the problems of semantic offsets and category boundary blur in short text clustering are solved, and a more efficient clustering effect is achieved.
Patent Information
- Application Number
- CN202510260961.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-20
AI Technical Summary
Short text clustering faces the problems of semantic offset and category boundary blur, which is difficult to accurately handle by traditional methods, resulting in unsatisfactory clustering effect.
Using a method based on a large language model, short text is converted into vector representations through Sentence-BERT, preclustered using HDBSCAN, and fuzzy samples between classes are sampled based on information entropy, and pseudo-labels are obtained using a large language model, combining data augmentation and instance-level comparison learning to achieve clustering.
It effectively solves the problems of semantic offset and category boundary blur, improves the effect of short text clustering, and can better deal with samples with category boundary blur.
Smart Images

Figure CN120179826A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and particularly to a short text clustering method, device, equipment and storage medium assisted by a large language model. Background Art
[0002] With the rapid development of information technology, digital content has shown an explosive growth. Among numerous information carriers, short texts have become one of the main forms of information transmission in today's network environment due to their concise and intuitive characteristics. The need for automatic grouping and classification of short texts has become increasingly prominent, which has promoted the development of short text clustering technology. This technology plays an important role in multiple application scenarios such as user intention understanding, personalized recommendation, and opinion analysis.
[0003] However, short text clustering faces unique challenges. Due to the common characteristics of short texts such as few words and casual expressions, traditional clustering methods relying on word frequency statistics often have difficulty accurately grasping the semantic content of texts. Such methods are prone to the problem of sparse feature matrices, resulting in unsatisfactory clustering effects.
[0004] In recent years, pre-trained language models represented by BERT have made breakthrough progress in the field of natural language processing, providing a new technical path for short text clustering. Currently, mainstream deep learning clustering methods usually use pre-trained models to extract text features and then input them into clustering algorithms. However, this method has two main problems: firstly, the feature extraction of pre-trained models is not optimized specifically for clustering tasks; secondly, in practical applications, short texts of different categories often overlap in the feature space, reducing the effect of clustering algorithms based on distance metrics. In addition, existing technical solutions have obvious deficiencies in dealing with samples with fuzzy category boundaries. Such samples usually possess the characteristics of multiple categories simultaneously and are difficult to accurately divide. Summary of the Invention
[0005] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a short text clustering method, device, equipment and storage medium assisted by a large language model. The method first uses Sentence-BERT to convert short texts into vector representations and performs pre-clustering through the HDBSCAN model; then samples inter-class fuzzy samples based on information entropy and uses the large language model to judge the cluster membership of these samples to obtain pseudo-labels; then performs data augmentation by means of synonym replacement, random deletion, word order adjustment, etc. to generate positive examples for each sample; finally, uses the pseudo-labels as additional positive examples and jointly participates in instance-level contrastive learning with the positive examples generated by data augmentation, and realizes clustering by combining category-level self-supervised representation learning. This method does not require preprocessing of data, can effectively process samples with fuzzy category boundaries, and improves the short text clustering effect. The present invention also provides corresponding devices, electronic devices and storage media.
[0006] A short text clustering method assisted by a large language model according to the present invention is carried out according to the following steps:
[0007] a. First, input the short text data set into the Sentence-Bert model to convert it into word vectors, and then use the word vectors for pre-clustering in the HDBSAN clustering model;
[0008] b. Sample inter-class fuzzy samples based on information entropy, and use the large language model to determine the cluster to which each inter-class fuzzy sample belongs as a pseudo-label; the inter-class fuzzy samples are samples at the class boundary that are difficult to accurately classify due to overlapping features with multiple classes;
[0009] The information entropy sampling method is as follows: First, calculate the probability that each sample belongs to each cluster center through the auxiliary target distribution, sort to obtain the two cluster centers with the highest probability of belonging to each sample as candidates, calculate the information entropy using the normalized probabilities of the candidate clusters and sort, and finally sample the top N samples;
[0010] c. Generate positive examples for each sample in the way of data augmentation to prepare positive examples for the next contrastive learning; the data augmentation methods include: synonym replacement, random word deletion, and word order swapping;
[0011] d. Use the pseudo-labels obtained in step b as additional positive examples, fuse the positive examples in step c for instance-level contrastive learning, and simultaneously jointly implement class-level self-supervised representation learning.
[0012] In step b, when using the large language model to determine the cluster to which each inter-class fuzzy sample belongs, first construct a prompt word, and with the help of the semantic understanding ability of the large language model, determine the cluster center with more similar semantics.
[0013] The specific operations of instance-level contrastive learning and class-level self-supervised learning in step d are as follows:
[0014] Use the pseudo-labels marked in step b as additional positive examples, fuse the positive examples obtained in step c, improve the loss function of traditional contrastive learning that only contains one positive example to make it have multiple positive examples; use other examples in the same batch as in-batch negative samples to implement contrastive learning with multiple positive examples;
[0015] Input the text vectors obtained from the short text data into the K-means clustering model to obtain the initial cluster centers; calculate the probability that each sample belongs to different cluster centers; enhance the label information of high-confidence labels by squaring the probability distribution; perform representation learning from high-confidence samples.
[0016] The device involved in the short text clustering method assisted by a large language model is composed of a sampling module based on information entropy, a large language model discrimination module, a category-level learning module, and an instance-level learning module, where:
[0017] Sampling module based on information entropy: used to sample fuzzy samples between categories;
[0018] Large language model discrimination module: used to perform pseudo-label annotation on fuzzy samples between categories;
[0019] Data augmentation module: used to generate positive examples through data augmentation methods;
[0020] Category-level learning module: used to provide category-level self-supervised representation learning for short text data;
[0021] Instance-level learning module: used to fuse multiple positive examples and perform contrastive representation learning at the instance level.
[0022] An electronic device for short text clustering assisted by a large language model, which includes: at least one processor, and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-3.
[0023] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to any one of claims 1-3.
[0024] The short text clustering method, device, equipment, and storage medium according to the present invention have the following technical effects:
[0025] Compared with the prior art, the present invention effectively solves problems such as semantic drift and fuzzy category boundaries in traditional short text clustering by integrating the semantic understanding ability of a large language model and a multi-positive example contrast learning strategy, improves the clustering effect, and can be widely applied to fields such as text analysis and intent recognition. Description of the Drawings
[0026] Figure 1 is the flowchart of the present invention;
[0027] Figure 2 is the flowchart of step 2 in the short text clustering method assisted by a large language model according to the present invention;
[0028] Figure 3 is the flowchart of step S4.2 in the short text clustering method assisted by a large language model according to the present invention;
[0029] Figure 4 This is a schematic structural diagram of an apparatus for short text clustering assisted by a large language model according to the present invention. Specific embodiments
[0030] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0031] Example 1
[0032] A method for short text clustering assisted by a large language model according to the present invention is carried out according to the following steps:
[0033] a. First, input the short text dataset into the Sentence-Bert model to convert it into word vectors, and then pre-cluster the word vectors using the HDBSAN clustering model.
[0034] b. Sample inter-class fuzzy samples based on information entropy, and use the large language model to determine the cluster to which each inter-class fuzzy sample belongs as a pseudo-label; the inter-class fuzzy samples are samples located at the class boundary and are difficult to accurately classify because their features overlap with multiple classes.
[0035] The information entropy sampling method is as follows: First, calculate the probability of each sample belonging to each cluster center through the auxiliary target distribution, sort to obtain the two cluster centers with the highest probability for each sample as candidates, calculate the information entropy using the normalized probabilities of the candidate clusters and sort, and finally sample the top N samples.
[0036] c. Generate positive examples for each sample using data augmentation to prepare positive examples for the next contrastive learning; the data augmentation methods include: synonym replacement, random word deletion, and swapping word order.
[0037] d. Use the pseudo-labels obtained in step b as additional positive examples, fuse the positive examples in step c for instance-level contrastive learning, and simultaneously jointly implement class-level self-supervised representation learning.
[0038] In step b, when using the large language model to determine the cluster to which each inter-class fuzzy sample belongs, first construct a prompt word, and with the help of the semantic understanding ability of the large language model, determine the cluster center with more similar semantics.
[0039] The specific operations of instance-level contrastive learning and class-level self-supervised learning in step d are as follows:
[0040] Take the pseudo-labels marked in step b as additional positive examples, fuse the positive examples obtained in step c, and improve the loss function of traditional contrast learning that only contains one positive example to make it have multiple positive examples; take other examples in the same batch as in-batch negative samples and implement contrast learning with multiple positive examples;
[0041] Input the text vectors obtained from the short text data into the K-means clustering model to obtain the initial cluster centers; calculate the probability that each sample belongs to different cluster centers; enhance the label information with high confidence by squaring the probability distribution; perform representation learning from the samples with high confidence.
[0042] Refer to Figure 1 , Figure 1 Figure is a schematic flowchart of a short text clustering method assisted by a large language model provided by the present invention. By innovatively combining the advantages of pre-trained models and large language models, an efficient short text clustering solution is provided, which is carried out according to the following steps:
[0043] Input the text vectors obtained from the short text data into the K-means clustering model to obtain the initial cluster centers; calculate the probability that each sample belongs to different cluster centers; enhance the label information with high confidence by squaring the probability distribution; perform representation learning from the samples with high confidence;
[0044] Step 1: First, input the short text data into Sentence-Bert to convert it into word vectors, and then use a clustering model without specifying the number of categories for pre-clustering;
[0045] This step aims to obtain a rough class division using the original embedding representation. The specific operation is as follows:
[0046] S1.1: Short text vectorization:
[0047] Input the short text to be clustered into the pre-trained Sentence-BERT model, and each short text is converted into a 768-dimensional vector representation;
[0048] S1.2: Short text pre-clustering:
[0049] Input the obtained vector representation into the HDBSCAN clustering model for preliminary clustering. The advantage of the HDBSCAN model is that it does not require specifying the number of clusters in advance and can automatically determine the number of categories according to the data distribution;
[0050] Step 2: Sample inter-class fuzzy samples based on information entropy, and use the large language model to determine the cluster to which each inter-class fuzzy sample belongs as a pseudo-label;
[0051] This step aims to sample the inter-class fuzzy samples in the dataset, such as Figure 2The specific operation is as follows:
[0052] S2.1: Probability distribution calculation:
[0053] Calculate the probability distribution of each sample belonging to each cluster center. The specific calculation formula for the probability that the i-th sample belongs to the c-th cluster is: where C is the number of clusters obtained by HDBSCAN clustering in S1.2; e i represents the vector obtained by the i-th sample through S1.1; μ c is the center vector of the c-th cluster; α is a parameter controlling the distribution shape;
[0054] S2.2: Select high-probability candidate clusters:
[0055] Sort the probabilities of each sample belonging to the cluster center and select the two cluster centers with the highest probabilities as candidates;
[0056] S2.3: Normalize probabilities and calculate information entropy:
[0057] Calculate the information entropy of the sample using the normalized probabilities of the candidate clusters. The formula for the normalized probability that the i-th sample belongs to the c-th cluster is where n = 2, which are the two clusters with the highest probabilities selected in S2.2; The calculation formula for the information entropy of the i-th sample is as
[0058] S2.4: Sample inter-class fuzzy samples:
[0059] Sort according to the information entropy from high to low and select the top N samples as inter-class fuzzy samples;
[0060] S2.5: Construct prompt templates:
[0061] For each inter-class fuzzy sample, construct a prompt template. For example: Among the following candidate categories, which category is the input text {{fuzzy sample}} semantically closest to? Candidate categories: {{Candidate category 1::center1_text}}, {{Candidate category 2:center2_text}};
[0062] S2.6: Judgment by large language model:
[0063] Fill the fuzzy sample and the representative text of its candidate clusters into the template, input it into the large language model, and obtain the category attribution judged by the model as the pseudo-label of the fuzzy sample;
[0064] Step 3: Generate positive examples for each sample using data augmentation to prepare positive examples for the next contrastive learning;
[0065] Specifically, for each sample in the dataset, the following data augmentation methods are used to generate positive samples:
[0066] Replace some words in the original text with words from a thesaurus of near synonyms; randomly delete non-keywords in the text; randomly swap the order of words; etc.
[0067] Step 4: Use the pseudo-labels obtained in Step 2 as additional positive examples, fuse the positive examples in Step 3, and perform instance-level contrastive learning, while jointly performing category-level self-supervised representation learning;
[0068] S4.1 Instance-level contrastive learning:
[0069] Use the positive samples generated in Step 3 as the first group of positive samples z α1 , z α2 , and use the other samples in the same training batch as negative examples whose vector representation is z n , and its contrastive learning loss is where τ is the temperature coefficient; sim(u, v) represents the cosine similarity between u and v;
[0070] Use the pseudo-labels discriminated by the large language model in Step 2 as additional positive samples z p , and use the other samples in the same training batch as negative examples whose vector representation is z n , and its contrastive learning loss is: where τ is the temperature coefficient; sim(u, v) represents the cosine similarity between u and v;
[0071] Jointly perform instance-level contrastive learning on the two groups of positive samples, and the overall loss of the final instance-level contrastive learning is where ε is a balancing parameter;
[0072] S4.2 Category-level self-supervised learning:
[0073] This step aims to perform category-level representation learning based on the current clustering results, optimize the semantic consistency within categories, and increase the distinguishability between categories, as Figure 3 shown, and its specific operation is:
[0074] S4.2.1 Obtain the initial cluster centers:
[0075] Input the text vectors obtained from the short text data into the K-means clustering model to obtain the initial cluster centers;
[0076] S4.2.2 Calculate the probability distribution:
[0077] Calculate the probability Q = {q ik | i = 1, 2, ..., N; k = 1, 2, ..., K} that each sample belongs to different cluster centers. The specific calculation formula for the i-th sample belonging to the k-th cluster is where K is the number of clusters; e i represents the vector representation of the i-th sample; μ k is the center vector of the k-th cluster; α is a parameter controlling the distribution shape;
[0078] S4.2.3 Soft Auxiliary Target Distribution:
[0079] Enhance the high-confidence label information by squaring the probability distribution, and then normalize it through the relevant clustering probabilities to obtain the soft auxiliary target P = {p ik | i = 1, 2, ..., N; k = 1, 2, ..., K}, and its calculation method is
[0080] S4.2.4 Class-Level Self-Supervised Learning:
[0081] Perform representation learning from high-confidence samples, and use the KL-divergence as the loss function to optimize the model. The calculation formula is as where Q is the probability distribution calculated in S4.2.2, and P is the soft auxiliary target distribution obtained in S4.2.3;
[0082] Finally, perform joint multi-objective optimization iterative learning, and the overall optimization objective is as shown in Equation where is the contrastive learning loss given in S4.1, is the class-level and self-supervised learning loss, is the overall loss.
[0083] Embodiment 2
[0084] The device involved in the short text clustering method assisted by the large language model is composed of a sampling module based on information entropy, a large language model discrimination module, a class-level learning module, and an instance-level learning module, where:
[0085] Sampling Module Based on Information Entropy: Used to sample fuzzy samples between classes;
[0086] Large Language Model Discrimination Module: Used to perform pseudo-label annotation on fuzzy samples between classes;
[0087] Data Augmentation Module: Used to generate positive examples through data augmentation methods;
[0088] Class-Level Learning Module: Used to provide class-level self-supervised representation learning for short text data;
[0089] Instance-level learning module: used to fuse multiple positive examples and perform instance-level contrastive representation learning;
[0090] As Figure 4 shown, a device for short text clustering assisted by a large language model is provided, which includes:
[0091] Information entropy-based sampling module: used to sample fuzzy samples between categories, which includes the following parts:
[0092] Text embedding generation part: using a pre-trained language model as an encoder to convert short texts into high-dimensional semantic embedding vectors, providing a basic representation for cluster analysis; Large language model discrimination module: used to perform pseudo-label annotation on fuzzy samples between categories;
[0093] Pre-clustering part: using a clustering algorithm to perform preliminary grouping on text embeddings, calculating the similarity probability distribution of samples with each cluster center, and providing a basis for uncertainty calculation;
[0094] Information entropy calculation part: calculating the uncertainty information entropy of each sample according to the probability distribution; the higher the entropy value, the more ambiguous the class attribution of the sample;
[0095] Fuzzy sample screening part: sorting the samples according to the entropy value and selecting samples with higher entropy values as fuzzy samples between categories, providing key data input for subsequent processing;
[0096] Candidate cluster screening part: screening out several cluster centers with the highest probability corresponding to each fuzzy sample to construct a candidate cluster set, reducing the complexity of sample processing, which includes the following functional parts:
[0097] Prompt template generation part: constructing a prompt template, splicing the fuzzy sample and its candidate cluster center text in a specific format to form a sample pair that meets the input requirements of the large language model;
[0098] Semantic matching judgment part: calling the large language model to process the prompt template, judging the semantic similarity between the fuzzy sample and the candidate cluster center, and generating a similarity score for each sample;
[0099] Pseudo-label generation part: according to the discrimination result, selecting the candidate cluster center with the closest semantics as the pseudo-label of the fuzzy sample and performing pseudo-label annotation on the sample to enhance the quality of clustering data;
[0100] Data augmentation module: used to generate positive examples through data augmentation methods, which includes parts for data augmentation such as synonym replacement, random word deletion, and swapping word order;
[0101] Category-level learning module: used to provide category-level self-supervised representation learning for short text data, which includes the following parts;
[0102] Intra - cluster consistency optimization part: Based on the probability distribution of the cluster center, adjust the similarity between samples within the cluster to make the samples within the same cluster more consistent in semantic representation;
[0103] Inter - cluster difference optimization part: Through the optimization of the auxiliary target distribution and KL divergence, enhance the distinguishability between different clusters and avoid excessive overlap of sample clusters;
[0104] Dynamic cluster distribution adjustment part: Combine fuzzy samples and pseudo - label annotation information to dynamically optimize the cluster distribution of each round of clustering, so that short texts gradually form a clearer semantic hierarchical structure in the representation space;
[0105] Instance - level learning module: Used to fuse multiple positive examples and perform instance - level contrastive representation learning, including the following functional parts:
[0106] Positive and negative example pair construction part: According to the fuzzy samples and their pseudo - label information, construct positive example pairs (fuzzy samples and their cluster centers) and negative examples (the remaining samples within the batch) for the construction of the contrastive learning objective function;
[0107] Contrastive learning optimization part: Through the contrastive learning framework, optimize the similarity between positive examples, and at the same time maximize the representation distance between positive and negative examples to improve the model's ability to distinguish different semantic categories;
[0108] Sample representation adjustment part: According to the contrastive learning objective function, adjust the distribution of samples in the representation space to make samples of the same category aggregated and samples of different categories separated, and optimize the clustering results;
[0109] Through the division of labor and cooperation of each module, this device realizes the complete process from the precise screening and annotation of fuzzy samples to data augmentation and deep optimization. Each functional part complements each other and jointly improves the accuracy and robustness of short - text clustering, providing an efficient and reliable technical solution for the semantic clustering task;
[0110] And an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any method described in the first aspect above;
[0111] Embodiment 3
[0112] An electronic device for short text clustering assisted by a large language model, including: at least one processor, and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-3.
[0113] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to any one of claims 1-3;
[0114] Specifically, the content of these instructions relates to various methods of the present invention, including but not limited to text vectorization, performing clustering calculations, short text data augmentation, and large language model invocation; when the instructions in the storage medium are executed by a computer, the method of short text clustering assisted by a large language model provided by the present invention can be implemented. Through this storage medium, the method of the present invention can actually run in a computer system and exhibit high flexibility and portability;
[0115] In the several embodiments provided, it should be understood that the devices and methods disclosed by the present invention can be implemented in various ways. The device and method embodiments described above are only illustrative, and the flowcharts in the drawings show the possible architectures, functions, and operations according to multiple embodiments of the present invention. In this regard, each block in the flowchart may represent a module, a program segment, or a part of code, which contains one or more executable instructions for implementing the specified logical function. At the same time, it should be noted that in alternative embodiments, the functions marked in the blocks may occur in a different order from that marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and sometimes may be executed in the reverse order, depending on the functions involved;
[0116] At the same time, the combination of each block in the block flowchart can be implemented by a dedicated hardware-based system that executes the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. Generally speaking, the embodiments of the present invention are not limited to the structures and steps shown in the drawings and the specific descriptions, but also cover various alternative embodiments and equivalent structures;
[0117] In addition, the functional modules in each embodiment can be integrated together to form an independent part, or each module can exist separately, and even two or more modules can be integrated to form an independent part;
[0118] When the functions are implemented in the form of software function modules and sold or used as independent products, these functions can be stored in a computer-readable storage medium. The part that essentially contributes to the technical solution of the present invention or in the prior art, as well as the part of the technical solution, can be presented in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (such as a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present invention. The storage medium can include various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
Claims
1. A short text clustering method based on a large language model, characterized in that The method proceeds in the following steps: a. First, input the short text dataset into the Sentence-Bert model and convert it into word vectors. Then, the word vectors are pre-clustered using the HDBSAN clustering model. b. Sampling inter-class fuzzy samples based on information entropy, and using a large language model to determine the cluster to which each inter-class fuzzy sample belongs as a pseudo label; the inter-class fuzzy samples are samples that are at the class boundary and are difficult to accurately classify because their features overlap with multiple categories; The information entropy-based sampling method is as follows: first, the probability of each sample belonging to each cluster center is calculated through the auxiliary target distribution, and the two cluster centers with the largest probability of each sample belonging are sorted as candidates, and the information entropy is calculated and sorted using the normalized probability of the candidate clusters, and finally the first N samples are sampled; c. Generate positive examples for each sample using data enhancement to prepare positive examples for the next step of comparative learning; the data enhancement methods include: synonym replacement, random word deletion, and word order exchange; d. Use the pseudo-labels obtained in step b as additional positive examples, integrate the instance-level contrastive learning of the positive examples in step c, and jointly implement category-level self-supervised representation learning.
2. A short text clustering method based on a large language model according to claim 1, characterized in that In step b, the large language model is used to determine the cluster to which each inter-class fuzzy sample belongs. First, a prompt word is constructed, and the semantic understanding ability of the large language model is used to determine the cluster center with more similar semantics.
3. A short text clustering method based on large language model assistance according to claim 1, characterized in that In step d, instance-level contrastive learning and category-level self-supervised learning are performed as follows: The pseudo-labels marked in step b are used as additional positive examples, and the positive examples obtained in step c are integrated to improve the loss function of traditional contrastive learning that only contains one positive example, so that it has multiple positive examples; other examples in the same batch are used as negative samples in the batch, and contrastive learning of multiple positive examples is implemented; Input the text vector obtained from the short text data into the K-means clustering model to obtain the initial cluster center; Calculate the probability that each sample belongs to the center of different clusters; The probability distribution is squared to enhance the high-confidence label information; representation learning is performed from high-confidence samples.
4. A device for the short text clustering method based on a large language model as claimed in claim 1, characterized in that: The device is composed of a sampling module based on information entropy, a large language model discrimination module, a category-level learning module and an instance-level learning module, wherein: Sampling module based on information entropy: used to sample fuzzy samples between categories; Large language model discrimination module: used to annotate pseudo-labels for ambiguous samples between categories; Data enhancement module: used to generate positive samples through data enhancement methods; Category-level learning module: used to provide category-level self-supervised representation learning for short text data; Instance-level learning module: used to fuse multiple positive examples and implement contrastive representation learning at the instance level.
5. An electronic device for short text clustering based on a large language model, wherein: include: At least one processor, and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 3.
6. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 3.
Citation Information
Cited By
Short text clustering and fuzzy recognition algorithm for large-scale network online sub-graph sampling
CN121958560A