A dialogue intent clustering method and system based on large language model cycle integration
Patent Information
- Application Number
- CN202610902273.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-25
AI Technical Summary
[0012]本发明的目的在于提供一种基于大语言模型循环集成的对话意图聚类方法及系统,以解决上述背景技术中提出的目前过度依赖嵌入距离度量、LLM集成深度不够和缺乏自动化的聚类数量发现机制等问题
第一,显著提升聚类质量。在自建中文客服数据集上,归一化互信息(NMI)达到0.8826,较K-means基线提升11.76%;语义一致性评分达到97.6%。
Smart Images

Figure CN122817469A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing technology, specifically to a method and system for clustering dialogue intent based on cyclic integration of large language models. Background Technology
[0002] Intent discovery is a key task in the field of natural language processing, widely used in dialogue system design, information retrieval, and discourse pattern analysis. Intent clustering technology helps build intelligent customer service systems and optimize user experience by automatically identifying topics and semantic structures in text corpora.
[0003] Currently, the mainstream intent clustering methods mainly include the following categories: The first category is clustering methods based on traditional machine learning, such as K-means, hierarchical clustering, and Gaussian mixture models. These methods rely on distance metrics (such as cosine similarity and Euclidean distance) of sentence embeddings for optimization. Although computationally efficient, they often ignore the semantic diversity and language pattern differences of text, making it difficult to accurately capture subtle differences in user intent, especially when dealing with sentences that have diverse expressions but the same intent.
[0004] The second category is representation learning methods based on deep learning, which obtain stronger semantic representations of sentences by training deep neural networks (such as BERT and Sentence-BERT). This type of method improves clustering results to some extent, but still mainly relies on distance metrics in the embedding space, and is prone to producing incorrect clustering when faced with sentences that are semantically similar but have different intentions.
[0005] The third category comprises LLM-guided clustering methods that have emerged in recent years, such as ClusterLLM, IDAS, and keyword expansion methods. These methods utilize large language models for data preprocessing, embedding optimization, or data augmentation. However, existing LLM-guided methods typically only use LLM for surface-level data integration, failing to deeply integrate its semantic understanding capabilities into the iterative optimization process of clustering, thus limiting further improvements in clustering performance.
[0006] In particular, in customer service dialogue scenarios, existing technologies face the following prominent issues: High semantic diversity: The same intent can be expressed in multiple ways, and embedding distance makes it difficult to accurately capture semantic consistency. For example, "What is universal life insurance?" and "Can you explain this universal life insurance to me in more detail?" express the same intent, but their cosine similarity is extremely low.
[0007] Difficulty in distinguishing intents: Different intents may use similar expressions, making them easy to confuse. For example, "Is it paying ten yuan in insurance premiums every month?" and "Is it deducting ten yuan from phone bills every month?" express different intents, but their cosine similarity is high.
[0008] Unknown number of clusters: In practical applications, it is usually impossible to know the true number of intent categories in advance. Traditional methods rely on heuristic rules such as the elbow rule, which are not ideal.
[0009] Evaluation metrics are inconsistent with human perception: Traditional metrics such as normalized mutual information (NMI) tend to favor a larger number of clusters, which deviates from human intuitive perception of cluster quality.
[0010] Insufficient dataset size and complexity: Existing English benchmark datasets are small in size and have limited semantic diversity. There is a lack of large-scale, highly complex Chinese customer service dialogue datasets, making it difficult to verify the effectiveness of the method in real-world scenarios.
[0011] Therefore, how to deeply integrate the semantic understanding capabilities of large language models into the iterative process of intent clustering, realize semantic-driven adaptive clustering, ensure that the clustering results are highly aligned with human perception, and improve the interpretability of clustering are technical problems that urgently need to be solved in this field. Summary of the Invention
[0012] The purpose of this invention is to provide a dialogue intent clustering method and system based on large language model cyclic integration, so as to solve the problems mentioned in the background art, such as over-reliance on embedding distance metrics, insufficient LLM integration depth, and lack of automated cluster number discovery mechanism.
[0013] To achieve the above objectives, the present invention provides the following technical solution: A dialogue intent clustering method based on cyclic ensemble of large language models includes the following steps: Step S1: Obtain the set of dialogue sentences to be clustered; Step S2: Map the dialogue sentences into semantic vectors using a sentence embedding model; Step S3: In each iteration, perform hierarchical clustering with multiple candidate cluster numbers on the currently unassigned sentences to generate multiple candidate cluster sets; Step S4: Use the fine-tuned large language model as a semantic consistency evaluator to perform a binary classification semantic consistency judgment of "Good" or "Bad" for each candidate cluster; Step S5: Select the optimal number of clusters based on the proportion of "Good" clusters in each candidate cluster set, retain the "Good" clusters with the optimal number as high-quality intent clusters in this round, and redistribute the sentences in the "Bad" clusters to the next iteration; Step S6: Repeat the above iterative process until the proportion of unassigned sentences is lower than the preset threshold or the maximum number of iterations is reached, and output the final high-quality intent cluster set.
[0014] As a further preferred technical solution, the training and usage process of the semantic consistency estimator includes: Construct training and validation sets for intent clusters manually labeled as "Good" or "Bad"; The LoRA method was used to fine-tune the Chinese large language model so that its output only contains the labels "Good" or "Bad". During the clustering process, representative sentences are sampled from each cluster and input into the fine-tuned large language model to obtain their semantic consistency labels.
[0015] As a further preferred technical solution, the sampling of representative sentences adopts a convex hull sampling strategy, including: Calculate the convex hull of the semantic vectors of all sentences within the cluster in the embedding space; The sentences corresponding to the convex hull vertices are selected as representative samples for input into the semantic consistency evaluator.
[0016] As a further preferred technical solution, a post-processing optimization step based on intent tags is also included: A finely tuned large language model is used as an intent label generator to generate semantic labels in "action-target" format for each high-quality intent cluster; The semantic tags are mapped to a unit hypersphere using a sentence embedding model, and the geodesic distance between the tags in the clusters is calculated. Model the uncertainty of label embedding based on the von Mises-Fisher distribution, and calculate the probability that any two clusters belong to the same intention; When the probability exceeds a preset threshold, the corresponding clusters are merged into a new intent cluster.
[0017] As a further preferred technical solution, the "action-goal" format intent tag generator is constructed in the following way: Collect sentence clusters from dialogues across multiple domains and manually label them with "action-goal" format tags; The LoRA method is used to fine-tune the Chinese large language model so that its output conforms to the "action-target" naming convention.
[0018] As a further preferred technical solution, a context-aware role separation step is also included: Based on the action portion of the intent labels in each cluster after clustering, heuristic rules are used to divide sentences into customer role groups and customer service role groups; The iterative clustering method described in claim 1 is re-executed for the customer role group and the customer service role group, respectively. By merging the two clustering results, the final set of role separation intention clusters is obtained.
[0019] As a further preferred technical solution, when selecting the optimal number of clusters in each iteration, a local search strategy that maximizes the ratio of the number of "Good" clusters to the number of "Bad" clusters is adopted, and different numbers of clusters are allowed to be selected in different iterations.
[0020] As a further preferred technical solution, the method is applied to Chinese customer service dialogue scenarios, where the set of dialogue sentences to be clustered consists of transcribed texts of Chinese customer service dialogues in the banking, telecommunications, or insurance sectors, with a sentence size of more than 50,000 sentences and more than 500 true intent categories, and the number of true categories does not need to be specified in advance during the clustering process.
[0021] A dialogue intent clustering system based on cyclic ensemble of large language models includes: The sentence embedding module is used to encode dialogue sentences into semantic vectors; The LLM semantic consistency evaluator is a fine-tuned large language model used to determine whether a family of sentences has a unified semantic intent. The LLM intent tag generator is a fine-tuned large language model used to generate "action-target" format tags for semantically consistent clusters. An iterative clustering engine for executing the iterative clustering process as described in claim 1; The post-processing module is used to perform the cluster merging operation as described in claim 4 and the role separation operation as described in claim 6.
[0022] As a further preferred technical solution, both the LLM semantic consistency evaluator and the LLM intent label generator are based on the Qwen2.5-7B, Qwen2.5-14B, Baichuan2-7B or ChatGLM3-6B models, and are fine-tuned on a manually annotated customer service dialogue intent cluster dataset using the LoRA method.
[0023] Compared with the prior art, the beneficial effects of the present invention are: First, it significantly improves clustering quality. On the self-built Chinese customer service dataset, the normalized mutual information (NMI) reaches 0.8826, an improvement of 11.76% compared to the K-means baseline; the semantic consistency score reaches 97.6%.
[0024] Second, it automatically discovers the number of clusters. Without needing to pre-specify the K value, the optimal number of clusters is automatically determined through semantic evaluation during the iterative process, adapting to situations where the intent category is unknown in real-world scenarios.
[0025] Third, the evaluation results are consistent with human judgment. The fine-tuned LLM evaluator achieved 97.5% consistency with human judgment, overcoming the shortcomings of traditional indicators that tend to favor more clusters.
[0026] Fourth, the clustering results are interpretable. Semantic tags in "action-target" format are generated for each intent cluster, facilitating manual review and downstream use.
[0027] Fifth, it has high computational efficiency. By using convex hull sampling, each cluster only requires 20 sentences of LLM input, and the number of calls is far fewer than similar methods.
[0028] Sixth, it fills a gap in Chinese datasets. A Chinese customer service dialogue dataset of 55,000 sentences and 1,507 intent categories was constructed, providing a realistic and complex benchmark for Chinese intent clustering. Attached Figure Description
[0029] Figure 1 This is a flowchart of a dialogue intent clustering method based on cyclic integration of a large language model according to the present invention. Figure 2 This is a system architecture diagram of a dialogue intent clustering system based on cyclic integration of a large language model according to the present invention. Detailed Implementation
[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] Please see Figure 1-2 This invention provides a technical solution: It proposes a dialogue intent clustering method and system based on LLM-in-the-loop (LLM-in-the-loop ensemble), which achieves high-quality, semantically driven intent discovery by deeply integrating a large language model into the iterative clustering process. The overall technical solution includes: LLM tool design, iterative clustering framework, post-processing optimization, and context-aware technology.
[0032] 1. System Architecture
[0033] The system of this invention includes the following core modules: (1) Sentence embedding module: Use a pre-trained dense encoder (such as BGE-large-zh) to encode dialogue sentences into high-dimensional semantic vectors.
[0034] (2) LLM Semantic Consistency Evaluator: A finely tuned large language model used to evaluate the semantic consistency of clusters and determine whether a cluster contains sentences with the same intent.
[0035] (3) LLM Intent Tag Generator: A finely tuned large language model used to generate semantic tags in "action-target" format for each intent cluster.
[0036] (4) Iterative clustering engine: Based on feedback from traditional clustering algorithms (such as hierarchical clustering) and LLM evaluators, iteratively discover high-quality intent clusters.
[0037] (5) Post-processing module: including cluster merging based on intent tags and context-aware separation based on role.
[0038] 2. LLM tool design
[0039] 2.1 Semantic Consistency Evaluator
[0040] Task Definition: Given a family of sentences To determine whether the cluster has semantic consistency, that is, whether all sentences express the same intent.
[0041] Labeling method: Use binary classification labeling, and label the clusters as "Good" (semantic consistency) or "Bad" (semantic inconsistency).
[0042] Training data construction: 1772 manually annotated intent clusters were collected as the training set; Good clusters: All sentences focus on the same theme and intention; Bad clusters: contain inconsistent or ambiguous intentions; The validation set consisted of 480 clusters used for evaluation; Model selection and fine-tuning: Basic models: Chinese LLMs such as Qwen2.5-7B, Qwen2.5-14B, Baichuan2-7B, and ChatGLM3-6B; Fine-tuning method: LoRA (low-rank adaptation); Hardware configuration: 4×Nvidia A100 GPUs; Performance: The finely tuned Qwen2.5-14B achieved 97.50% accuracy on the validation set; Input format: Sample representative sentences from each cluster (typical sample size: 20 sentences), the input format is as follows: You are a sentence clustering aid. Based on the relevance and commonalities of the following sentence clusters, classify them as either "Good" or "Bad". Only provide labels; do not output any additional content.
[0043] Example: Input: [List of sentences] Output: [Labels]
[0044] Input: {[list of sentences]} Output: Output format: Single label "Good" or "Bad"
[0045] 2.2 Intent Tag Generator
[0046] Task definition: Given a semantically consistent set of sentences, generate concise intent labels using the "action-goal" naming convention.
[0047] Naming convention design: The "action-goal" format is particularly suitable for capturing dialogue intent, because dialogue intent is usually topic-oriented (such as insurance, loans) and involves specific actions (such as asking questions, confirming).
[0048] Example: Inquiry - Universal Life Insurance: Users are asking about information related to universal life insurance. Answer - Amount: Customer service answers questions about the specific amount. Confirm - Identity: Verify user identity information Training data construction: 2500 intent clusters, each containing 20 sentences; Manually label intent tags in "action-goal" format; It covers three major sectors: banking, telecommunications, and insurance. Model fine-tuning: Uses the same basic model as the consistency evaluator; After fine-tuning, the accuracy rates were 94.3% for Qwen 2.5-14B and 94.4% for ChatGLM 3-6B. Input format: You are a sentence clustering aid. Summarize the following sentence clusters using "action-goal" labels based on their relevance and commonalities. Provide only the labels; do not output any additional content.
[0049] Example: Input: [List of sentences] Output: [Labels]
[0050] Input: {[list of sentences]} Output: Output format: Single "Action-Target" format tag
[0051] 3. Data sampling strategy
[0052] To reduce LLM call costs while maintaining evaluation accuracy, this invention designs two sampling strategies to select representative sentences from each cluster: (1) Random Sampling: Randomly select n sentences from the cluster (typically n=20). Simple and efficient, but may miss boundary samples. (2) Convex hull sampling: Calculate the convex hull of the cluster in the sentence embedding space. The sentence corresponding to the vertex of the convex hull is selected as the representative. It can better cover the semantic boundaries of the cluster. Experiments show that convex hull sampling is superior to random sampling, improving accuracy by approximately 2-3 percentage points.
[0053] 4. LLM-in-the-loop iterative clustering algorithm
[0054] The core of this invention is to deeply integrate LLM semantic evaluation into the clustering iteration process to automatically discover the optimal number of clusters.
[0055] 4.1 Algorithm Flow
[0056] enter: S: The set of unlabeled sentences f_emb: Sentence embedding function (e.g., BGE-large-zh) M_eval: Semantic Consistency Evaluator (Fine-tuned LLM) N: The set of candidate cluster sizes, such as N={10,30,50,70,90,110} ε: Stop threshold (e.g., ε=0.05 means stop when 5% of the sentences remain) T_max: Maximum number of iterations (e.g., T_max=10) Output: A set of high-quality intent clusters C Algorithm steps: Algorithm 1: LLM-in-the-loop Intended Clustering
[0057] (1) Initialization
[0058] S^(0)←S / / Set of unassigned sentences in round 0
[0059] C← / / Final cluster set
[0060] t←0 / / Iteration counter
[0061] (2) Iterative clustering
[0062] while|S^(t)| / |S|>εandt <T_max:
[0063] 1) Embedding of currently unassigned sentences
[0064] E^(t)←{f_emb(s)|s∈S^(t)}
[0065] 2) Cluster the number of each candidate cluster.
[0066] foreach n_i∈N:
[0067] 3) Use LLM to evaluate the semantic consistency of each cluster.
[0068] g^(t)_{n_i}←[ ]
[0069] foreachclusterC_j∈C^(t)_{n_i}:
[0070] sample_j←ConvexSa l_j←M_ Returns "Good" or "Bad"
[0071] g^(t)_{n_i}.ap
[0072] 4) Select the optimal number of clusters (maximizing the good / bad ratio)
[0073] 5) Retain the "Good" cluster and remove the assigned sentences.
[0074] C^(t)_good←{C_j∈C^(t)_{n }|g^(t)_{n }[j]==1}
[0075] C←C∪C^(t)_good
[0076] S^(t+1)←S^(t)\∪_{C∈C^(t)_good}C
[0077] t←t+1
[0078] (3) Return results
[0079] returnC
[0080] 4.2 Key Technologies
[0081] (1) Adaptive clustering quantity discovery: The optimal K value is automatically selected by evaluating the good / bad ratio under different cluster sizes. Each iteration may select a different number of clusters to adapt to changes in data distribution. This avoids the limitation of traditional methods that require pre-specifying K. (2) Iterative optimization strategy: In each iteration, only the "Good" cluster is retained, while sentences in the "Bad" cluster are left for re-clustering in the next round. As the iterations continue, the number of unassigned sentences gradually decreases, reducing the clustering difficulty. The process stops when the proportion of unassigned sentences falls below a threshold ε. The remaining sentences can be labeled as noise or out-of-domain data. (3) Clustering algorithm selection: Basic clustering algorithm: Hierarchical clustering Distance metric: cosine distance Linking method: Average Linkage Experiments show that hierarchical clustering outperforms K-means (1.29% improvement in NMI) and GMM (1.24% improvement in NMI) in this task.
[0082] 5. Post-processing optimization based on intent tags
[0083] Iterative clustering may produce multiple small clusters that express similar intentions, which need to be merged to improve clustering quality and interpretability.
[0084] 5.1 Intent Tag Generation
[0085] For each cluster Ck, generate labels using the LLM intent label generator: Return to the "Action-Target" format tab
[0086] 5.2 Tag Embedding and Hyperspherical Mapping
[0087] Map intent tags to sentence embedding space:
[0088] Normalized to a unit hypersphere :
[0089] Advantages of hypersphere mapping: On a unit hypersphere, semantic relationships can be accurately measured using geodesic distance, rather than straight-line distance in Euclidean space.
[0090] 5.3 Semantic Affinity Graph Construction
[0091] Construct an inter-cluster semantic affinity graph G=(V,E): vertex Represents all clusters Edge E is determined based on the geodesic distance of the label embedding. Geodesic distance calculation (angular distance on the hypersphere):
[0092] in This indicates the inner product.
[0093] Determining the edges: if (Typical value θ = 0.8 radians, approximately 45.8 degrees), then candidate edges are established between Ci and Cj.
[0094] 5.4 Probabilistic Merging Decision
[0095] To improve the robustness of the merger decision, the von Mises-Fisher distribution is used to model the uncertainty of label embedding.
[0096] von Mises-Fisher distribution:
[0097] in: It is the mean direction of the m-th intention. >0 controls the concentration of the distribution (typical value = 10) It is a normalization constant It is a modified Bessel function of the first kind. Agreement probability calculation:
[0098] in =1 / K represents the uniform mixing weight.
[0099] Merge conditions: Only when (When the typical value τ=0.7) edges are preserved And perform the merge.
[0100] 5.5 Connected Component Merging
[0101] Use graph connectivity algorithms (such as depth-first search (DFS) or union-find) to find all connected components, and merge the clusters in each connected component into a new cluster:
[0102] Intent labels are regenerated for each merged cluster to form the final clustering result.
[0103] Merging effect: Experiments show that this method can improve NMI from 0.8208 to 0.8420, an improvement of 2.58%.
[0104] 6. Context-aware role separation technology
[0105] Customer service conversations typically involve two roles: the customer and the customer service representative. The intentions of different roles have different characteristics, and mixed clustering may degrade quality. This invention proposes an unsupervised role separation method based on intent labeling.
[0106] 6.1 Role Recognition Based on Intent Labels
[0107] Core idea: Infer the role of a sentence by examining the "action" part of the intent tag.
[0108] Heuristic rules: Clusters of sentences containing actions such as "inquiry," "consultation," "complaint," and "request," originating from customers. Clusters of sentences containing actions such as "answer -", "confirm -", "process -", and "provide -", sourced from customer service. Role grouping: R_customer={s∈S|l_k(s)∈{"Inquiry-","Consultation-","Complaint-",...}} R_agent={s∈S|l_k(s)∈{"Answer-","Confirm-","Process-",...}} Where l_k(s) represents the intent label of the cluster to which sentence s belongs.
[0109] 6.2 Two-stage clustering process
[0110] Phase 1: Perform preliminary clustering on all sentences to obtain intermediate results C_inter and intent labels.
[0111] Phase Two: 1. Based on intent tags, sentences are divided into customer group R_customer and customer service group R_agent. 2. Re-cluster the two groups of sentences separately: ... C'_customer=LLM-ITL-Cluste te Combine the clustering results of the two groups: ... C_final=C'_customer∪C'_agent ... Advantages: To avoid misclustering of similar expressions from different characters Improve the semantic purity of clustering Experiments show that role separation can improve NMI from 0.8208 to 0.8679, an improvement of 5.74%.
[0112] 7. Construction of a large-scale Chinese customer service dialogue dataset
[0113] To verify the effectiveness of the method, this invention constructed the largest Chinese customer service dialogue intent clustering dataset to date.
[0114] 7.1 Data Collection
[0115] Data source: Transcriptions of over 100,000 real customer service phone calls, covering the banking, telecommunications, and insurance sectors.
[0116] Data scale: Original conversation: 11,879 calls Filtered conversations: 8,184 (sensitive information removed) Total number of sentences: 69,839 Number of sentences after deduplication: 55,085 (unique sentences) Average sentence length: 17 Chinese characters
[0117] 7.2 Data Labeling
[0118] Labeling Team: 15 professional labelers, each with more than 5 years of experience in the customer service field.
[0119] Labeling process: 4. Use K-means initial clustering (K=2000) as the starting point. 5. The annotator evaluates the semantic consistency (Good / Bad) of each cluster. 6. Label the Good cluster with intent tags in "action-goal" format. 7. Reassign sentences in the Bad cluster or create new clusters. 8. Repeat steps 2-4 until all sentences are correctly categorized. Annotation results: Total intent categories: 1,507 Domain-specific intent: 885 (related to banking, telecommunications, and insurance) Out-of-domain intents: 622 (general queries, such as providing location, confirmation time, etc.) Semantic diversity: 0.538 (significantly higher than the 0.209-0.367 of existing English datasets) Quality control: Verification was conducted by 10 independent experts. Resolving labeling disagreements through consensus. Ensure the reliability and consistency of the dataset.
[0120] 7.3 Dataset Characteristics
[0121] (1) Large scale: 55,085 sentences and 1,507 intent categories, which is more than 10 times the existing largest intent clustering dataset.
[0122] (2) High semantic diversity: multiple ways of expressing the same intention and similar expressions of different intentions fully reflect the complexity of real-world scenarios.
[0123] (3) Realistic noise: It contains a large number of foreign queries, colloquial expressions, and incomplete sentences, which are close to actual applications.
[0124] (4) Chinese characteristics: fully embodying the semantic richness and expressive diversity of the Chinese language.
[0125] Advantages of this solution compared to existing technologies: (1) The clustering quality has been significantly improved: On a self-built Chinese dataset, the NMI reached 0.8826, an improvement of 11.76% compared to the K-means baseline (0.7899). The Goodness score reached 97.6%, an improvement of 2.95% compared to the K-means baseline (94.8%). On English benchmark datasets, its performance is comparable to or better than state-of-the-art methods (NMI reaches 78.12% on the MASSIVE dataset). (2) Automatically discover the optimal number of clusters: No need to pre-specify the number of clusters K; clusters are automatically discovered through an iterative process. Each iteration can select a different value for K to adapt to changes in data distribution. It avoids the uncertainty inherent in traditional methods that rely on heuristic rules (such as the elbow rule). (3) The evaluation indicators are consistent with human perception: The finely tuned LLM evaluator achieved 97.5% consistency with human judgment. The Goodness metric directly reflects the semantic consistency of clustering, overcoming the bias of NMI towards more clusters. The generated "action-target" labels are intuitive and easy to understand, facilitating manual review and downstream applications. (4) Excellent performance in downstream applications: The clustered data generated using this method was used to train an intent classifier, achieving an accuracy of 77%. An improvement of 18.46% compared to the K-means baseline (65%). It also has significant advantages compared to other LLM bootstrapping methods (ClusterLLM 62%, IDAS 68%). (5) High computational efficiency: By employing a convex hull sampling strategy, each cluster only needs to sample 20 sentences of input LLM. Compared to the pairwise comparisons of ClusterLLM (which requires 1024 LLM calls), this method requires only about 560 calls on the Bank77 dataset. Supports offline caching of LLM judgment results to avoid repeated calls during training. (6) High robustness: By using probabilistic merging decision-making (von Mises-Fisher distribution), the oversensitivity of deterministic methods is avoided. Role separation technology effectively handles the unique structure of customer service dialogues, improving NMI by 5.74%. It performs stably on real-world data containing significant noise and out-of-domain queries. (7) Cross-language generalization ability: The accuracy rate on the Chinese dataset reached 97.6% (Goodness). Its performance on English datasets is comparable to state-of-the-art methods. The framework design is language-independent and can be extended to other languages. (8) Good scalability: It can be flexibly integrated with different basic clustering algorithms (K-means, hierarchical clustering, GMM, etc.) Different LLM models (open source or commercial APIs) can be used. The sampling strategy and threshold parameters can be adjusted according to application requirements.
[0126] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A dialogue intent clustering method based on cyclic ensemble of large language models, characterized in that, Includes the following steps: Step S1: Obtain the set of dialogue sentences to be clustered; Step S2: Map the dialogue sentences into semantic vectors using a sentence embedding model; Step S3: In each iteration, perform hierarchical clustering with multiple candidate cluster numbers on the currently unassigned sentences to generate multiple candidate cluster sets; Step S4: Use the fine-tuned large language model as a semantic consistency evaluator to perform a binary classification semantic consistency judgment of "Good" or "Bad" for each candidate cluster; Step S5: Select the optimal number of clusters based on the proportion of "Good" clusters in each candidate cluster set, retain the "Good" clusters with the optimal number as high-quality intent clusters in this round, and redistribute the sentences in the "Bad" clusters to the next iteration; Step S6: Repeat the above iterative process until the proportion of unassigned sentences is lower than the preset threshold or the maximum number of iterations is reached, and output the final high-quality intent cluster set.
2. The method according to claim 1, characterized in that, The training and use process of the semantic consistency evaluator includes: Construct training and validation sets for intent clusters manually labeled as "Good" or "Bad"; The LoRA method was used to fine-tune the Chinese large language model so that its output only contains the labels "Good" or "Bad". During the clustering process, representative sentences are sampled from each cluster and input into the fine-tuned large language model to obtain their semantic consistency labels.
3. The method according to claim 1, characterized in that, The sampling of the representative sentences adopts a convex hull sampling strategy, including: Calculate the convex hull of the semantic vectors of all sentences within the cluster in the embedding space; The sentences corresponding to the convex hull vertices are selected as representative samples for input into the semantic consistency evaluator.
4. The method according to claim 1, characterized in that, It also includes post-processing optimization steps based on intent tags: A finely tuned large language model is used as an intent label generator to generate semantic labels in "action-target" format for each high-quality intent cluster; The semantic tags are mapped to a unit hypersphere using a sentence embedding model, and the geodesic distance between the tags in the clusters is calculated. Model the uncertainty of label embedding based on the von Mises-Fisher distribution, and calculate the probability that any two clusters belong to the same intention; When the probability exceeds a preset threshold, the corresponding clusters are merged into a new intent cluster.
5. The method according to claim 4, characterized in that, The intent tag generator in the "action-goal" format is constructed in the following way: Collect sentence clusters from dialogues across multiple domains and manually label them with "action-goal" format tags; The LoRA method is used to fine-tune the Chinese large language model so that its output conforms to the "action-target" naming convention.
6. The method according to claim 1, characterized in that, It also includes a context-aware role separation step: Based on the action portion of the intent labels in each cluster after clustering, heuristic rules are used to divide sentences into customer role groups and customer service role groups; The iterative clustering method described in claim 1 is re-executed for the customer role group and the customer service role group, respectively. By merging the two clustering results, the final set of role separation intention clusters is obtained.
7. The method according to claim 1, characterized in that, When selecting the optimal number of clusters in each iteration, a local search strategy is adopted to maximize the ratio of the number of "Good" clusters to the number of "Bad" clusters, and different numbers of clusters are allowed to be selected in different iterations.
8. The method according to claim 1, characterized in that, The method is applied to Chinese customer service dialogue scenarios, where the set of dialogue sentences to be clustered consists of transcribed texts of Chinese customer service dialogues in the banking, telecommunications, or insurance sectors, with a sentence size of more than 50,000 sentences and more than 500 true intent categories. Furthermore, the number of true categories does not need to be specified in advance during the clustering process.
9. A dialogue intent clustering system based on cyclic ensemble of large language models, characterized in that, include: The sentence embedding module is used to encode dialogue sentences into semantic vectors; The LLM semantic consistency evaluator is a fine-tuned large language model used to determine whether a family of sentences has a unified semantic intent. The LLM intent tag generator is a fine-tuned large language model used to generate "action-target" format tags for semantically consistent clusters. An iterative clustering engine for executing the iterative clustering process as described in claim 1; The post-processing module is used to perform the cluster merging operation as described in claim 4 and the role separation operation as described in claim 6.
10. The system according to claim 9, characterized in that, The LLM semantic consistency evaluator and LLM intent label generator are both based on the Qwen2.5-7B, Qwen2.5-14B, Baichuan2-7B or ChatGLM3-6B models, and are fine-tuned on a manually annotated customer service dialogue intent cluster dataset using the LoRA method.