Intelligent customer service system customer multi-level intent recognition method based on stage learning
Patent Information
- Application Number
- CN202511330205.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-09-17
AI Technical Summary
[0002]随着人工智能技术的迅速崛起,智能客服成为企业提升服务质量、降低服务成本的重要手段,在众多应用场景中,如何准确识别用户在对话中表达的多层次意图,是智能客服系统亟待解决的技术问题;现有技术中,通常采用基于深度学习的分类方法,在大量语料库上训练LLM模型后再用自己数据集进行调优来获得分类效果,然而,这种方法在处理简单意图时效果尚可,但对于智能客服系统中层次众多的意图识别就难以应对,容易出现意图理解偏差、分类边界模糊等问题
[0037]The beneficial effects of this invention are as follows: The multi-level intent recognition method for intelligent customer service systems based on phased learning provided by this invention can accurately identify multi-level intents in user dialogues under limited annotation cost constraints. It achieves complete intent understanding from coarse-grained to fine-grained through a hierarchical classification system, improving the accuracy and hierarchy of intent recognition. By constructing a three-layer intent classification system and adopting a hybrid supervised learning strategy, it fully utilizes fully labeled data, weakly supervised data, and unlabeled data, ensuring model performance while reducing annotation costs and maximizing the utilization of data resources. By using a weighted loss calculation method to assign differentiated penalty weights to different types of classification errors, it effectively avoids serious errors interfering with the intelligent customer service process. Using a hierarchical consistency loss function and a predefined hierarchical projection matrix, the probability distributions of the middle and bottom layers are aggregated upwards to reconstruct the top-level probability distribution, ensuring that the three-layer intent prediction results remain logically consistent and preventing contradictions between predictions at different levels. By employing a phased training strategy, the system focuses on building top-level intent recognition capabilities during the basic intent establishment phase. In the hierarchical structure self-discovery phase, it achieves autonomous discovery of hierarchical intents through pseudo-label generation and dialogue coherence constraints. In the fine-grained optimization phase, it optimizes the classification boundaries of lower-level intents through a contrastive learning task, thus realizing a progressive learning process from coarse to fine. By utilizing the dialogue coherence prediction task to uncover the inherent logical relationships between adjacent dialogue statements, combined with a clustering pseudo-label generation strategy and a hierarchical negative sample construction mechanism, the model's robustness and generalization ability in dialogue scenarios are ensured, effectively improving the service quality and user experience of the intelligent customer service system.
Smart Images

Figure CN121388090B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing technology, specifically to a method for recognizing multi-level customer intent in an intelligent customer service system based on phased learning. Background Technology
[0002] With the rapid rise of artificial intelligence technology, intelligent customer service has become an important means for enterprises to improve service quality and reduce service costs. In many application scenarios, how to accurately identify the multi-layered intentions expressed by users in the dialogue is a technical problem that intelligent customer service systems urgently need to solve. In existing technologies, classification methods based on deep learning are usually adopted. The LLM model is trained on a large corpus and then optimized with its own dataset to obtain the classification effect. However, this method is acceptable when dealing with simple intentions, but it is difficult to deal with the multi-layered intention recognition in intelligent customer service systems, and it is easy to have problems such as intention misunderstanding and blurred classification boundaries.
[0003] Furthermore, intelligent customer service systems have extremely stringent requirements for intent recognition, as incorrect recognition results directly impact user satisfaction and service experience. In practice, different types of recognition errors have varying degrees of severity. For example, misidentifying "balance inquiry" as "package inquiry" is a low-level error, while misidentifying "inquiry" as "complaint" is a serious error that crosses the top-level intent, causing the intelligent customer service response to enter an incorrect process. Secondly, existing end-to-end training methods treat all intents as equally important and train them together. This strategy struggles to handle hierarchical intent classification, especially when the amount of labeled data is limited, easily leading to uneven learning across layers and poor classification results. On the other hand... The development of intelligent customer service intent recognition technology is limited by the difficulty in constructing high-quality datasets. High-quality multi-level intent recognition datasets require a large amount of manual annotation, which is not only costly and time-consuming, but also requires annotators to have in-depth business understanding and professional knowledge to ensure annotation accuracy. In actual projects, only a small amount of fully annotated data can usually be obtained, while a large amount of dialogue data either has only coarse-grained annotations or no annotations at all. Although traditional semi-supervised learning methods can utilize unlabeled data to some extent, their effectiveness is limited in multi-level intent recognition tasks. Moreover, existing methods lack effective utilization of dialogue context information and ignore the coherence and logical characteristics of user intent in customer service dialogues. Summary of the Invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] To address the aforementioned technical problems, this invention provides the following technical solution: a multi-level customer intent recognition method for an intelligent customer service system based on phased learning, comprising:
[0006] Acquire fully labeled data, weakly supervised data, and unlabeled data. Based on the fully labeled data and the weakly supervised data, set a loss function that emphasizes the top-level intent and establish a basic intent recognition model.
[0007] Based on the basic intent recognition model, mid-level intent pseudo-labels and low-level intent pseudo-labels are generated for the weakly supervised data. Combined with dialogue coherence loss, the basic intent recognition model is trained using the unlabeled data to obtain a hierarchical intent recognition model.
[0008] Based on the hierarchical intent recognition model, the underlying intent is optimized through a comparative learning task to obtain a multi-level intent recognition model.
[0009] The dialogue data of user questions is obtained, and the multi-level intent recognition model is used to perform intent recognition on the dialogue data to obtain the recognition results of top-level intent, middle-level intent and bottom-level intent.
[0010] As a preferred embodiment of the multi-level customer intent recognition method for intelligent customer service system based on phased learning described in this invention, the following is a preferred embodiment: setting a loss function that emphasizes the top-level intent based on the fully labeled data and the weakly supervised data to establish a basic intent recognition model includes setting a first loss function for the fully labeled data that includes the top-level intent loss, the middle-level intent loss and the bottom-level intent loss, wherein the weight of the top-level intent loss is greater than the weight of the middle-level intent loss and the bottom-level intent loss.
[0011] A second loss function, incorporating the top-level intent loss, is applied to the weakly supervised data.
[0012] A basic intent recognition model is established by jointly training the first loss function and the second loss function.
[0013] As a preferred embodiment of the multi-level customer intent recognition method for an intelligent customer service system based on phased learning described in this invention, the method for generating mid-level intent pseudo-labels and bottom-level intent pseudo-labels for the weakly supervised data based on the basic intent recognition model includes: using the basic intent recognition model to perform top-level intent prediction on the weakly supervised data to obtain the top-level intent prediction confidence level; filtering weakly supervised data whose top-level intent prediction confidence level is higher than a preset threshold to obtain the filtered weakly supervised data samples.
[0014] Based on the real top-level intent labels of the selected weakly supervised data samples, random sampling is performed within the corresponding mid-level intent subclasses using a predefined hierarchical mapping relationship to generate mid-level intent pseudo-labels.
[0015] Based on the mid-level intent pseudo-tags, random sampling is performed within the corresponding low-level intent subclass range to generate low-level intent pseudo-tags.
[0016] As a preferred embodiment of the multi-level customer intent recognition method for an intelligent customer service system based on phased learning described in this invention, the method includes: training the basic intent recognition model using the unlabeled data in conjunction with dialogue coherence loss, which involves constructing a dialogue coherence prediction task and determining the intent coherence between adjacent dialogue statements.
[0017] The unlabeled data is processed using the dialogue coherence prediction task to calculate the dialogue coherence loss.
[0018] The mid-level intent pseudo-labels and the bottom-level intent pseudo-labels are combined with the dialogue coherence loss to train the basic intent recognition model, resulting in a hierarchical intent recognition model.
[0019] As a preferred embodiment of the multi-level customer intent recognition method for an intelligent customer service system based on phased learning described in this invention, wherein: based on the hierarchical intent recognition model, optimizing the underlying intent through a comparative learning task includes performing clustering processing on the feature representation output by the hierarchical intent recognition model to generate clustering pseudo-labels for the unlabeled data;
[0020] Construct positive sample pairs, which include sample pairs with the same real labels and sample pairs with the same cluster pseudo labels;
[0021] Construct negative sample pairs, which include sample pairs of different intent categories;
[0022] The hierarchical intent recognition model is trained using the positive and negative sample pairs to optimize the classification boundary of the underlying intent.
[0023] As a preferred embodiment of the multi-level customer intent recognition method for the intelligent customer service system based on phased learning described in this invention, wherein: clustering the feature representation output by the hierarchical intent recognition model and generating cluster pseudo-labels for the unlabeled data includes extracting the feature vectors output by the feature sharing layer of the three classification heads in the hierarchical intent recognition model as clustering objects;
[0024] The feature vectors are clustered using a clustering algorithm, with the number of clusters set to the total number of categories for the top-level intent, middle-level intent, and bottom-level intent.
[0025] The clustering results are tracked over multiple consecutive training cycles, and pseudo-labels are assigned to samples whose clustering results are consistent and whose samples are close to the cluster center.
[0026] As a preferred embodiment of the multi-level customer intent recognition method for an intelligent customer service system based on phased learning as described in this invention, the construction of negative sample pairs includes constructing a first type of negative sample pair across the top-level intent category;
[0027] Construct a second type of negative sample pair that belongs to the same top-level intent but to different mid-level intents;
[0028] Based on the prediction confidence of the hierarchical intent recognition model, samples with high confidence but incorrect predictions are mined to construct a third type of negative sample pair;
[0029] The first type of negative sample pairs, the second type of negative sample pairs, and the third type of negative sample pairs are combined to form hierarchical negative sample pairs.
[0030] A multi-level customer intent recognition system based on phased learning, wherein:
[0031] The basic intent establishment module acquires fully labeled data, weakly supervised data, and unlabeled data. Based on the fully labeled data and the weakly supervised data, it sets a loss function that focuses on the top-level intent and establishes a basic intent recognition model.
[0032] The hierarchical structure self-discovery module generates mid-level and low-level pseudo-labels for the weakly supervised data based on the basic intent recognition model, and trains the basic intent recognition model using the unlabeled data in conjunction with dialogue coherence loss to obtain a hierarchical intent recognition model.
[0033] The fine-grained optimization module, based on the hierarchical intent recognition model, optimizes the underlying intent through a comparative learning task to obtain a multi-level intent recognition model;
[0034] The intent recognition module acquires dialogue data from user questions and uses the multi-level intent recognition model to perform intent recognition on the dialogue data, obtaining recognition results for top-level intent, middle-level intent, and bottom-level intent.
[0035] A computer device includes: a memory and a processor; the memory stores a computer program, wherein: when the processor executes the computer program, it implements the steps of the method described in any one of the present invention.
[0036] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of the present invention.
[0037] The beneficial effects of this invention are as follows: The multi-level intent recognition method for intelligent customer service systems based on phased learning provided by this invention can accurately identify multi-level intents in user dialogues under limited annotation cost constraints. It achieves complete intent understanding from coarse-grained to fine-grained through a hierarchical classification system, improving the accuracy and hierarchy of intent recognition. By constructing a three-layer intent classification system and adopting a hybrid supervised learning strategy, it fully utilizes fully labeled data, weakly supervised data, and unlabeled data, ensuring model performance while reducing annotation costs and maximizing the utilization of data resources. By using a weighted loss calculation method to assign differentiated penalty weights to different types of classification errors, it effectively avoids serious errors interfering with the intelligent customer service process. Using a hierarchical consistency loss function and a predefined hierarchical projection matrix, the probability distributions of the middle and bottom layers are aggregated upwards to reconstruct the top-level probability distribution, ensuring that the three-layer intent prediction results remain logically consistent and preventing contradictions between predictions at different levels. By employing a phased training strategy, the system focuses on building top-level intent recognition capabilities during the basic intent establishment phase. In the hierarchical structure self-discovery phase, it achieves autonomous discovery of hierarchical intents through pseudo-label generation and dialogue coherence constraints. In the fine-grained optimization phase, it optimizes the classification boundaries of lower-level intents through a contrastive learning task, thus realizing a progressive learning process from coarse to fine. By utilizing the dialogue coherence prediction task to uncover the inherent logical relationships between adjacent dialogue statements, combined with a clustering pseudo-label generation strategy and a hierarchical negative sample construction mechanism, the model's robustness and generalization ability in dialogue scenarios are ensured, effectively improving the service quality and user experience of the intelligent customer service system. Attached Figure Description
[0038] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 The flowchart illustrates the overall process of a multi-level customer intent recognition method for an intelligent customer service system based on phased learning, as provided in the first embodiment of the present invention.
[0040] Figure 2 A diagram illustrating multi-layered consciousness provided for the first embodiment of the present invention.
[0041] Figure 3 This is a structural diagram of the LLM model provided in the first embodiment of the present invention.
[0042] Figure 4 This is a schematic diagram of a phased training process provided in the first embodiment of the present invention. Detailed Implementation
[0043] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0044] Example 1, referring to Figures 1-4 As an embodiment of the present invention, a method for recognizing multi-level customer intent in an intelligent customer service system based on phased learning is provided, comprising:
[0045] S1: Obtain fully labeled data, weakly supervised data, and unlabeled data. Based on the fully labeled data and the weakly supervised data, set a loss function that emphasizes the top-level intent and establish a basic intent recognition model.
[0046] In this embodiment, a clear hierarchical intent is represented by constructing three layers of intent, specifically as follows: Figure 2 As shown, the three layers of intent are K1, K2 and K3, where K1 is the top layer with a intents; K2 is the middle layer with b intents; and K3 is the finest classification layer, i.e. the bottom layer, with c intents. Therefore, the model has a total of a+b+c intent classification outputs, where c>a>b.
[0047] Secondly, addressing the issues of inaccurate multi-layered user intent recognition and high dataset labeling costs in existing intelligent customer service systems, a data utilization strategy based on hybrid supervised learning is employed. The acquisition of fully labeled data, weakly supervised data, and unlabeled data refers to constructing the following dataset to effectively utilize the labeled information: Seed dataset: D1 fully labeled data entries (i.e., K1, K2, and K3 layers are all labeled); Weakly supervised data: D2 data entries with only K1 labeled entries; Unlabeled data: D3 raw dialogue data entries (D1... <D2<D3)。
[0048] In real-world projects, due to the cost constraints of annotation, only a small amount of fully annotated data can be obtained. Meanwhile, a large amount of dialogue data either has only coarse-grained annotations (such as only annotating the top-level intent) or no annotations at all. By mixing a large amount of unlabeled data to form the dataset used for training, the problem of high-quality multi-level intent recognition datasets requiring a large amount of manual annotation is effectively alleviated. This is not only costly and time-consuming, but also requires annotators to have a deep understanding of the business.
[0049] Since relying solely on a single LLM model (BERT, Llama3, Qwen3) is insufficient for effectively extracting semantic feature vectors and accurately classifying intents, it is essential to use different classification heads to classify intents at different levels. Therefore, this embodiment designs a method as follows: Figure 3 The model structure is shown below. Specifically, a pre-trained LLM model (BERT, Llama3, Qwen3) is used to extract semantic features from the input dialogue. The obtained semantic feature vector is then processed through a 3-layer MLP and three classifiers are used to output the intent probability of each layer. The MLP uses ReLU activation and Dropout regularization. The three classifiers share weights, and finally, the output vector is projected into three dimensions (a, b, c) using a Softmax function. By sharing the underlying feature extraction, it is ensured that the intent classification at the three levels is based on a consistent semantic understanding. At the same time, by using independent classifiers, the model can be specifically optimized for intent classification tasks of different granularities, improving the model's ability to recognize hierarchical intents.
[0050] The requirements for intent recognition in intelligent customer service systems are quite stringent, as incorrect recognition results directly lead to user dissatisfaction and experience problems. In practice, different types of recognition errors have different degrees of severity. For example, misidentifying "balance inquiry" as "package inquiry" is a relatively low-level error; however, misidentifying "inquiry" as "complaint" is a serious error that crosses the top-level intent, causing the intelligent customer service response to enter another process. Therefore, this embodiment proposes a weighted loss calculation method, which guides the model to learn the hierarchical relationship between intents by assigning different penalty weights to different types of classification errors (e.g., errors that cross the top-level intent have the highest weight). When intent recognition is incorrect, it is generally believed that recognition errors across parent classes are more serious than recognition errors within child classes. Therefore, K1 layer loss, K2 layer loss, and K3 layer loss are introduced, as follows:
[0051] The K1 layer loss (i.e., semantically weighted distance cross-entropy) ensures that the model has a robust classification ability for top-level intents by employing standard cross-entropy loss, laying the foundation for subsequent hierarchical intent recognition. The specific formula is as follows:
[0052]
[0053] in, The loss is K1 layer, i.e., semantically weighted distance cross-entropy; N is the total number of samples in the current batch; y ij Let i be the true label of the i-th sample for the j-th category; Let be the predicted probability of the i-th sample for the j-th K1 category; a is the total number of intent categories in the K1 layer.
[0054] The K2 layer loss (i.e., hierarchical perceptual cross-entropy) introduces hierarchical weights, which more severely penalize classification errors across K1 categories, thus reflecting the importance of parent-child category relationships in hierarchical intent recognition. The specific formula is as follows:
[0055]
[0056] in, The loss is K2 layer, i.e., hierarchical perceptual cross-entropy; N is the total number of samples in the current batch; h ij For hierarchical weights, the classification error within the same K1 level is... Intra-K1 classification error and By establishing the indicator weights through fuzzy hierarchical analysis, we can ensure that the penalties for classification errors at different levels meet actual business needs. Let be the predicted probability of the i-th sample for the j-th K2 category; b is the total number of intent categories in the K2 layer.
[0057] The K3 layer loss (i.e., the three-level hierarchical classification cross-entropy) uses a hierarchical weight design to differentiate the penalties for classification errors at different levels, ensuring that the model prioritizes the correctness of higher-level intentions. The specific formula is as follows:
[0058]
[0059] in, The loss is the K3 layer loss, i.e., the three-level hierarchical classification cross-entropy; N is the total number of samples in the current batch; g ij The classification weights are divided into three levels: ω1 for classification errors within the same K2 level, ω2 for classification errors across K2 levels but correct classification results in K1 level, and ω3 for classification errors in K1 level (ω3>ω2>ω1). ω1, ω2, and ω3 are determined by the index weights of fuzzy hierarchical analysis to ensure that the penalty for classification errors at different levels meets the actual business needs. Let be the predicted probability of the i-th sample for the j-th K3 category; c is the total number of intent categories in the K3 layer.
[0060] Furthermore, to ensure the logical consistency of the three-layer prediction results and prevent logically flawed predictions, a hierarchical consistency loss function is constructed. This function uses a predefined hierarchical projection matrix to aggregate the probability distributions of K2 and K3 upwards, reconstructing the probability distribution of K1. A KL divergence penalty is then applied to distributions with significant prediction discrepancies, ensuring logical consistency in the model's three-layer intent prediction results. Thus, even without fully labeled three-layer information, the model can discover the hierarchical structure of intent during training. The specific formula for the hierarchical consistency loss function is as follows:
[0061]
[0062] in, The hierarchical consistency loss function; The consistency loss is calculated from K2 to K1. The consistency loss is calculated from K3 to K2. The consistency loss is from K1 to K2 and K3.
[0063] Using KL divergence to assess the consistency loss from K1 to K2 and K3 Reconstruction is performed to obtain the KL divergence loss from K1 to K2 and K3 after reconstruction. The reconstruction function uses the hierarchical projection matrix, and the specific formula is as follows:
[0064]
[0065] in, This refers to the KL divergence loss from K1 to K2 and K3 after reconstruction; and Let N be the probability distribution of the i-th sample in layers K1, K2, and K3; N is the total number of samples in the current batch; Reconstruct is the reconstruction function used to reconstruct the upper probability distribution from the lower probability distribution; Projection is the projection aggregation matrix method.
[0066] To enable the model to better learn hierarchical classification boundaries on data with limited fully labeled information, weakly labeled data, and unlabeled data, this embodiment employs a three-stage training strategy, as detailed below. Figure 4 As shown, the training process includes a basic intent establishment phase (the first 30% of training epochs), a hierarchical structure self-discovery phase (the middle 30% of training epochs), and a fine-grained optimization phase (the last 40% of training epochs). Different types of labeled data and loss functions are used in combination at different training phases, enabling the model to complete different tasks at different stages and thus more effectively distinguish the classification boundaries.
[0067] In the foundational intent building phase (the first 30% of training epochs), the goal of this embodiment is to establish a robust K1 layer classification capability, enabling the model to achieve high classification accuracy at the K1 layer and laying a solid foundation for subsequent training. During this phase, each training batch uses 80% fully labeled data as the primary data and 20% weakly labeled data as auxiliary data.
[0068] Furthermore, based on the fully labeled data and the weakly supervised data, a loss function emphasizing top-level intent is set to establish the basic intent recognition model. This includes setting a first loss function for the fully labeled data that includes top-level intent loss, mid-level intent loss, and bottom-level intent loss, wherein the weight of the top-level intent loss is greater than the weights of the mid-level and bottom-level intent losses. A second loss function that includes the top-level intent loss is set for the weakly supervised data. The basic intent recognition model is then established through joint training using the first loss function and the second loss function.
[0069] The first loss function for the fully labeled data includes top-level intent loss, mid-level intent loss, and bottom-level intent loss. For fully labeled data, by setting the K1 layer loss to dominate, the model is ensured to focus on learning the classification ability of top-level intents in the first stage, while maintaining a basic understanding of mid-level and bottom-level intents. The specific formula of the first loss function is as follows:
[0070]
[0071] in, The first loss function; The loss is the K1 layer loss, which is the semantically weighted distance cross-entropy. The loss is the K2 layer loss, i.e., the hierarchical perceptual cross-entropy. The loss is the K3 layer loss, which is the three-level hierarchical classification cross-entropy. This is the hierarchical consistency loss function.
[0072] Setting a second loss function that includes the top-level intent loss for the weakly supervised data means maximizing the value of the weakly supervised data by focusing on the calculation of the K1 layer loss, thus providing support for the establishment of basic intent recognition capabilities. Since the weakly supervised data only contains K1 layer annotation information, only the K1 layer loss is calculated. The specific formula for the second loss function is as follows:
[0073]
[0074] in, This is the second loss function.
[0075] The basic intent recognition model is established by jointly training using the first loss function and the second loss function. This involves jointly training on fully labeled data and weakly supervised data, which fully utilizes the rich information in the fully labeled data and expands the scale of the training data, thereby improving the model's generalization ability. The overall loss function is obtained, and the specific formula is as follows:
[0076]
[0077] in, This is the overall loss function.
[0078] In addition, during the training process of establishing the basic intent, a layer-by-layer unfreezing strategy is used to update the parameters by gradually unfreezing the pre-trained LLM and the three classifiers. When this stage of training is completed, all parameters should be unfrozen, and different learning rates should be used to update the parameters for different layers of classifiers. For example, a larger learning rate can be set for the classifiers of layer K1, and the learning rates of the classifiers of layers K2 and K3 can be decreased successively. This layered learning rate method is also used in subsequent stages.
[0079] To address the issues of inaccurate multi-level user intent recognition and high dataset labeling costs in existing intelligent customer service intent recognition methods, this paper establishes a basic intent recognition model by constructing a three-layer intent classification system (K1, K2, and K3 layers) and employing a hybrid supervised learning data utilization strategy. Specifically, a three-branch neural network model based on a shared feature layer is designed. Semantic features are extracted using a pre-trained LLM and processed by a multilayer perceptron. Three independent classification heads then output intent prediction results at different granularities, ensuring consistent semantic understanding across the three layers. Furthermore, a weighted loss calculation method is introduced to assign differentiated penalty weights to different types of classification errors, with errors crossing the top-level intent receiving the highest weight, guiding the model to learn the hierarchical relationship between intents. In addition, a hierarchical consistency loss function is constructed, using a predefined hierarchical projection matrix to aggregate the probability distributions of the middle and bottom layers upwards to reconstruct the top-level probability distribution. KL divergence is used to penalize distributions with significant differences, ensuring logical consistency in the three-layer intent prediction results. This maximizes the value of different types of data while maintaining the quality of labeled data, laying a solid foundation for subsequent hierarchical intent learning.
[0080] S2: Based on the basic intent recognition model, generate mid-level intent pseudo-labels and low-level intent pseudo-labels for the weakly supervised data, and combine them with dialogue coherence loss. Then, train the basic intent recognition model using the unlabeled data to obtain a hierarchical intent recognition model.
[0081] During the self-discovery phase of the hierarchical structure (the middle 30% of the training epochs), the goal of this embodiment is to allow the model to spontaneously discover the hierarchical structure of intents, so that it can correctly identify child intents even if the parent intent is misidentified, thereby correcting the parent intent error. If constraints are imposed on the hierarchical structure in the data, i.e., if the model's classification of K1 is a... iThis forces the model to limit its intent recognition for layers K2 and K3 to the categories of layer K1. However, this would cause the model to continue making the wrong predictions for layer K1 intents. Therefore, the model needs to be allowed to discover the hierarchical structure of intents spontaneously. Thus, after the basic intent recognition model has learned the layer K1 intent classification, in this stage, pseudo-labels for layer K2 intents and layer K3 intents are generated using data with weak supervision information. This, combined with fully labeled data, allows the model to spontaneously discover hierarchical intents.
[0082] Furthermore, generating mid-level and low-level intent pseudo-labels for the weakly supervised data based on the basic intent recognition model includes: using the basic intent recognition model to perform top-level intent prediction on the weakly supervised data, obtaining a top-level intent prediction confidence score, and filtering weakly supervised data with a top-level intent prediction confidence score higher than a preset threshold to obtain filtered weakly supervised data samples. Based on the true top-level intent labels of the filtered weakly supervised data samples, random sampling is performed within the corresponding mid-level intent subclasses using a predefined hierarchical mapping relationship to generate mid-level intent pseudo-labels. Based on the mid-level intent pseudo-labels, random sampling is performed within the corresponding low-level intent subclasses to generate low-level intent pseudo-labels.
[0083] It should be noted that the top-level intent prediction confidence refers to the reasoning of the weakly supervised data by the basic intent recognition model obtained during the basic intent establishment period, to obtain the K1 layer intent prediction probability distribution of each sample, and to calculate the prediction confidence. The prediction confidence measures the model's confidence in the top-level intent prediction of the sample by taking the maximum value in the probability distribution. The higher the confidence, the more reliable the model's K1 layer prediction of the sample.
[0084] Filtering weakly supervised data with a top-level intent prediction confidence level higher than a preset threshold means that, in order to ensure the quality of pseudo-label generation, pseudo-labels are only generated for data with a K1 intent confidence level greater than a preset threshold. By setting a reasonable confidence threshold, samples with relatively reliable predictions at the K1 layer can be selected, avoiding the negative impact of pseudo-label noise caused by low-confidence predictions on subsequent training. This reflects the principle of quality over quantity, and it is better to reduce the scale of pseudo-label data than to ensure the reliability of pseudo-labels. In this example, the preset threshold is set to 90%.
[0085] By employing a progressive pseudo-label generation strategy, mid-level and low-level intention pseudo-labels are obtained. This approach takes into account the fact that, in the absence of precise K2 annotations, introducing appropriate randomness can enhance the model's generalization ability while avoiding overfitting to specific K2 categories. It also ensures the consistency of the hierarchical structure. Furthermore, random sampling provides the model with diverse training signals, which helps the model learn more robust hierarchical intention representations.
[0086] It should be noted that in customer service dialogues, users' intentions are usually consistent at the session level, that is, users will not arbitrarily jump between completely unrelated intentions in the same session. Therefore, this embodiment constructs a dialogue coherence prediction task, which mines the internal logical relationship of the dialogue by judging the intention coherence between adjacent dialogue sentences. The dialogue coherence prediction task takes two adjacent dialogue sentences as input and outputs a binary classification result of whether the dialogue is coherent in terms of intention.
[0087] Furthermore, training the basic intent recognition model using the unlabeled data, combined with the dialogue coherence loss, includes constructing a dialogue coherence prediction task to determine the intent coherence between adjacent dialogue statements. The dialogue coherence prediction task is used to process the unlabeled data to calculate the dialogue coherence loss. The mid-level intent pseudo-labels and the low-level intent pseudo-labels are combined with the dialogue coherence loss to train the basic intent recognition model, resulting in a hierarchical intent recognition model.
[0088] The unlabeled data is processed using the aforementioned dialogue coherence prediction task. The dialogue coherence loss is calculated using a binary cross-entropy loss function. By introducing this function, the model can learn the inherent patterns of intent changes in the dialogue. Even without complete three-layer annotation information, the learning of hierarchical intent can be guided by dialogue coherence constraints. The specific formula for the dialogue coherence loss function is as follows:
[0089]
[0090] p(consistent|u t-1 ,u t ))];
[0091] in, The dialogue coherence loss function is used; N is the total number of samples in the current batch; u t-1 For the preceding sentence (the text above); u t For the next sentence (below); y i For true and consistent labels; (1-y i ) represents inconsistent labels; logp(consistent|u t-1 ,u t This is used to predict the probability that the two sentences are semantically (intent-wise) consistent.
[0092] The mid-level and low-level intent pseudo-labels are combined with the dialogue coherence loss to train the basic intent recognition model, resulting in a hierarchical intent recognition model. This model includes a hierarchical self-discovery loss function to enable the model to spontaneously discover reasonable and accurate hierarchical intents, establishing a two-way constraint mechanism from fine-grained to coarse-grained levels. The specific formula is as follows:
[0093]
[0094] in, Hierarchical self-discovery loss function; CE is the cross-entropy loss function; W K2→K1 W is a predefined projection matrix from layer K2 to layer K1. K3→K2 This is a predefined projection matrix from layer K3 to layer K2; Let be the predicted probability distribution of the i-th sample in layer K2; λ is the weighting coefficient of the bottom-up constraint. Let be the predicted probability distribution of the i-th sample in layer K3.
[0095] Top-down constraints utilize a predefined projection matrix W K2→K1 The fine-grained intent (K2) distribution predicted by the model The distribution is aggregated into a coarse-grained intent (K1) and then compared with the true label using cross-entropy loss. Consistency is ensured, guaranteeing that fine-grained predictions do not deviate from the true semantics of coarse-grained predictions. Bottom-up constraints utilize a predefined projection matrix W. K3→K2 Distribute more fine-grained intentions (K3) The data is aggregated to a K2 layer and then subjected to cross-entropy loss to align it with the K2 distribution predicted by the model itself, i.e., the fine-grained intent (K2) distribution. Maintain consistency and ensure that the lowest-level predictions are semantically constrained by the predictions directly above them.
[0096] In this embodiment, when training the basic intent recognition model during the hierarchical self-discovery period, 40% of fully labeled data, 50% of weakly supervised data, and 10% of unlabeled data are used. By adjusting the proportion of different types of data, the value of weakly supervised data and unlabeled data is maximized while ensuring the stability of the model, and the dependence on fully labeled data is reduced.
[0097] Furthermore, during the hierarchical structure self-discovery phase, the loss weight ratio for pseudo-labeled data is reduced to 0.5 to mitigate the problem of inaccurate pseudo-labels. The main objective of this stage is to enable the model to self-discover the intended hierarchical structure; therefore, the three-layer classification losses should be relatively similar. The loss functions for fully labeled data, weakly supervised data, and unlabeled data, as well as the overall loss function, are as follows:
[0098] The specific formula for the loss function of fully labeled data is as follows:
[0099]
[0100] in, It is the loss function for fully labeled data during the self-discovery period of the hierarchical structure.
[0101] The specific formula for the loss function of weakly supervised data (including pseudo-labeled data) is as follows:
[0102]
[0103] in, This is the loss function for weakly supervised data during the self-discovery period of the hierarchical structure.
[0104] The specific formula for the loss function of unlabeled data is as follows:
[0105]
[0106] in, This is the loss function for unlabeled data during the self-discovery period of the hierarchical structure.
[0107] The specific formula for the overall loss function is as follows:
[0108]
[0109] in, This is the overall loss function during the self-discovery period of the hierarchical structure.
[0110] In this embodiment, by constructing loss functions for fully labeled data, weakly supervised data, and unlabeled data, as well as an overall loss function, the self-discovery learning of hierarchical intent is effectively achieved. On the one hand, the quality and hierarchical consistency of pseudo-labels are ensured through high-confidence screening and layer-by-layer pseudo-label generation strategies. On the other hand, by introducing dialogue coherence constraints, the inherent logical relationships in unlabeled data are used to guide model learning. At the same time, through the hierarchical self-discovery loss function, a two-way constraint mechanism from fine-grained to coarse-grained is established, enabling the model to autonomously discover the hierarchical structure of intent in the absence of complete annotations, thereby improving the model's ability and robustness in recognizing hierarchical intent.
[0111] To address the issues of uneven learning and recurring errors across layers when dealing with hierarchical intent classification using traditional training methods, this paper achieves self-discovery learning of hierarchical intents through a pseudo-label generation strategy based on the basic intent recognition model and dialogue coherence constraints. Specifically, by filtering the confidence level of top-level intent predictions on weakly supervised data, pseudo-labels are generated only for high-confidence samples. A progressively layered random sampling strategy is used to generate pseudo-labels for mid- and low-level intents, ensuring both the quality of pseudo-labels and enhancing the model's generalization ability. Furthermore, by constructing a dialogue coherence prediction task, the intent coherence between adjacent dialogue statements is utilized. By uncovering the inherent logical relationships within dialogues, the model can learn the patterns of intent changes within the dialogue. Secondly, by introducing a hierarchical self-discovery loss function, a two-way constraint mechanism is established from fine-grained to coarse-grained. The top-down constraint ensures that fine-grained predictions do not deviate from the true semantics of coarse-grained predictions, while the bottom-up constraint ensures that the lowest-level predictions are subject to the semantic constraints of the predictions directly above them. This enables the model to autonomously discover the hierarchical structure of intents even in the absence of complete annotations. Even if the parent intent is misidentified, it can still correctly identify the child intent and correct the parent's error, avoiding the error propagation problem caused by forced hierarchical constraints.
[0112] S3: Based on the hierarchical intent recognition model, the underlying intent is optimized through a comparative learning task to obtain a multi-level intent recognition model.
[0113] During the fine-grained optimization phase (the last 40% of training epochs), since the model has already achieved good classification of layer K1 and has a clear understanding of hierarchical intent in the first two phases, the main task of this embodiment is to make the model's classification of layer K3 more refined and accurate, and to improve the model's generalization ability by utilizing a large amount of unlabeled data. In the fine-grained optimization phase, contrastive learning is introduced as an auxiliary training method to improve the model's ability to judge classification boundaries.
[0114] Furthermore, based on the introduced contrastive learning, negative sample pairs are automatically generated through the hierarchical structure learned by the hierarchical intent recognition model, and positive sample pairs are constructed from data with labeled information. During the fine-grained optimization period, the model uses positive and negative sample pairs to bring dialogues with the same intent closer together in the vector space and push away dialogues with different intents, thereby becoming better able to distinguish classification boundaries. The specific formula for contrastive learning is as follows:
[0115]
[0116] in, For contrastive learning loss function; z i The feature representation of the anchor sample is extracted from the output of the feature sharing layer of the neural network and used as the center of the contrastive learning to compare with other samples; The feature representation of positive samples; zk is the feature representation of all samples in the batch; K is the batch size; sim(·) is used to calculate the semantic similarity between two feature vectors, and the most commonly used similarity calculation method is cosine similarity; τ is the temperature parameter. A smaller temperature parameter τ makes the model more sensitive to similarity differences and the clustering more compact, while a larger temperature parameter τ makes the distribution smoother and avoids overfitting.
[0117] Furthermore, based on the hierarchical intent recognition model, optimizing the underlying intent through a contrastive learning task includes: clustering the feature representation output by the hierarchical intent recognition model to generate cluster pseudo-labels for the unlabeled data; constructing positive sample pairs, which include sample pairs with the same real labels and sample pairs with the same cluster pseudo-labels; constructing negative sample pairs, which include sample pairs of different intent categories; and training the hierarchical intent recognition model using the positive and negative sample pairs to optimize the classification boundary of the underlying intent.
[0118] Specifically, clustering the feature representation output by the hierarchical intent recognition model to generate cluster pseudo-labels for the unlabeled data includes extracting the feature vectors output by the feature sharing layer of the three classification heads in the hierarchical intent recognition model as clustering objects. A clustering algorithm is used to cluster the feature vectors, with the number of clusters set to the total number of categories for top-level intents, middle-level intents, and bottom-level intents. Clustering results are tracked over multiple consecutive training periods, and cluster pseudo-labels are assigned to samples with consistent clustering results and close to the cluster centers.
[0119] It should be noted that, in order to enable a large amount of unlabeled data to participate in contrastive learning and improve the model's generalization ability, a cluster pseudo-label generation strategy is used during the fine-grained optimization period. The generation of cluster pseudo-labels can provide category information for unlabeled data, enabling it to participate in contrastive learning. By increasing the scale and diversity of training data, the model's ability to recognize fine-grained intents at the K3 layer is improved. Specifically, the K-means clustering algorithm is used to cluster the 256-dimensional feature vectors output by the feature sharing layer, and the number of clusters K is set to a+b+c (corresponding to the total number of intent categories). To ensure clustering quality, pseudo-labels are only assigned to the 50% of samples whose clustering results are consistent over three consecutive epochs and are closest to the cluster center. Boundary data does not participate in contrastive learning, thereby ensuring the stability and reliability of the cluster pseudo-labels and avoiding noise interference caused by cluster instability.
[0120] Furthermore, the construction of positive sample pairs refers to the identification of positive sample pairs following these rules:
[0121] For labeled data, samples with the same real label constitute positive sample pairs.
[0122] For unlabeled data, cluster pseudo-labels are used, and samples in the same cluster form positive sample pairs.
[0123] The construction of negative sample pairs includes: constructing a first type of negative sample pair that spans the top-level intent category; constructing a second type of negative sample pair that belongs to the same top-level intent but to different mid-level intents; and, based on the prediction confidence of the hierarchical intent recognition model, mining samples with high confidence but incorrect predictions to construct a third type of negative sample pair. The first, second, and third types of negative sample pairs are then combined to form hierarchical negative sample pairs.
[0124] Furthermore, in this embodiment, the construction of negative sample pairs adopts a hierarchical strategy, including three types of negative samples: simple, medium, and difficult, following the rules below:
[0125] Simple negative samples: For K2 and K3, random samples are taken across K1 to construct negative sample pairs across the top-level intent category; among them, simple negative samples are semantically different and are easily distinguished by the model, mainly used to establish the model's basic ability to distinguish between different top-level intents.
[0126] Medium negative samples: Construct negative sample pairs that belong to the same top-level intent but different mid-level intents, i.e. sample pairs of the same K1 category but different K2 subclasses; among them, medium negative samples are the same in top-level intents but have differences in mid-level intents, requiring the model to have more detailed discrimination capabilities.
[0127] Difficult sample mining: Difficult samples are identified in each epoch based on the model's prediction confidence. Samples with a prediction confidence higher than 0.8 but classified incorrectly are used as difficult negative samples for key training. Difficult samples are highly deceptive to the model. By learning from these difficult samples, the robustness of the model can be improved.
[0128] In this embodiment, when training the hierarchical intent recognition model during the fine-grained optimization period, 30% of fully labeled data, 30% of weakly supervised data, and 40% of unlabeled data are used. By increasing the proportion of unlabeled data, and with fully labeled data and weakly supervised data each accounting for half of the remaining data, the generalization ability of the model is maximized by utilizing unlabeled data.
[0129] Furthermore, the main task during the fine-grained optimization phase is to improve the model's ability to classify data with fine precision. Therefore, the classification loss of the K3 layer should account for a larger proportion. Since the complete labels for weakly supervised data are generated using pseudo-labels from the hierarchical self-discovery phase, the corresponding loss weight should be less than that for fully labeled data and unlabeled data using clustering pseudo-labels. For other losses, the weight should be greater than that from the hierarchical self-discovery phase. The loss functions for fully labeled data, weakly supervised data, and unlabeled data, as well as the overall loss function, are as follows:
[0130] The specific formula for the loss function of fully labeled data is as follows:
[0131]
[0132] in, This is the loss function for fully labeled data during fine-grained optimization.
[0133] The specific formula for the loss function of weakly supervised data is as follows:
[0134]
[0135] in, This is the loss function for weakly supervised data during fine-grained optimization.
[0136] The specific formula for the loss function of unlabeled data is as follows:
[0137]
[0138] in, This is the loss function for unlabeled data during fine-grained optimization.
[0139] The specific formula for the overall loss function is as follows:
[0140]
[0141] in, This is the overall loss function during the fine-grained optimization period.
[0142] To address the issue of insufficient ability to distinguish fine-grained intent classification boundaries in the model, K-means clustering was performed by extracting feature vectors from the feature-sharing layer of the hierarchical intent recognition model to generate high-quality cluster pseudo-labels for unlabeled data. A clustering result tracking and distance-to-cluster-center filtering mechanism was employed across multiple training cycles to ensure the stability and reliability of the cluster pseudo-labels. By constructing negative sample pairs with three difficulty levels—simple, medium, and hard—where simple negative samples span the top-level intent category, medium negative samples belong to the same top-level intent but different mid-level intents, and hard negative samples are obtained by mining high-confidence but incorrectly predicted samples, the model received progressive training from basic to advanced levels. Furthermore, by using a comparative learning loss function to shorten the distance between samples with the same intent and widen the distance between samples with different intents in the vector space, clear category boundaries were established, enabling the model to more accurately distinguish fine-grained intent differences. Overall, by increasing the proportion of unlabeled data, the potential value of the data was fully explored, improving the model's accuracy in recognizing low-level intents while enhancing its generalization and boundary judgment capabilities.
[0143] S4: Obtain the dialogue data of the user's question, and use the multi-level intent recognition model to perform intent recognition on the dialogue data to obtain the recognition results of the top-level intent, the middle-level intent, and the bottom-level intent.
[0144] The process of acquiring user-generated dialogue data and using the multi-level intent recognition model to identify the intent in the dialogue data to obtain the recognition results of top-level intent, middle-level intent, and bottom-level intent refers to acquiring the user-input dialogue text in real time from the intelligent customer service system, preprocessing the dialogue data, and then inputting it into the trained multi-level intent recognition model. The multi-level intent recognition model extracts semantic feature vectors from the dialogue data through a pre-trained LLM. After passing through three layers of MLP, the semantic feature vectors are input into three classification heads, which output the intent probability distributions of layers K1, K2, and K3, respectively.
[0145] For example, the final output recognition results include: K1 layer outputs top-level intent classification results, such as query, complaint, and processing; K2 layer outputs mid-level intent classification results, such as balance query and package query; K3 layer outputs bottom-level intent classification results, providing the finest-grained intent classification. Through hierarchical consistency constraints, the prediction results of the three layers are ensured to be logically consistent, providing accurate multi-level intent recognition results for the intelligent customer service system, thereby achieving more precise service response.
[0146] By extracting semantic features from user dialogues through pre-trained LLM, and then inputting them into three independent classification branches after feature transformation by a multilayer perceptron, the system simultaneously outputs intent recognition results at three granularities: top-level, middle-level, and bottom-level. This provides the intelligent customer service system with a complete intent understanding from coarse-grained to fine-grained. Through hierarchical consistency constraints, the system ensures that the prediction results of the three layers are logically consistent, avoiding contradictions between prediction results at different levels and improving the reliability of the recognition results. This enables the intelligent customer service system to accurately understand the user's specific needs and directly call the corresponding business process interfaces, avoiding potential intent understanding deviations or the need for multiple interactions for confirmation.
[0147] On the other hand, this embodiment also provides a multi-level customer intent recognition system for intelligent customer service systems based on phased learning, which includes:
[0148] The basic intent establishment module acquires fully labeled data, weakly supervised data, and unlabeled data. Based on the fully labeled data and the weakly supervised data, it sets a loss function that emphasizes the top-level intent and establishes a basic intent recognition model.
[0149] The hierarchical structure self-discovery module generates mid-level and low-level pseudo-labels for the weakly supervised data based on the basic intent recognition model, and trains the basic intent recognition model using the unlabeled data in conjunction with dialogue coherence loss to obtain a hierarchical intent recognition model.
[0150] The fine-grained optimization module, based on the hierarchical intent recognition model, optimizes the underlying intent through a comparative learning task to obtain a multi-level intent recognition model.
[0151] The intent recognition module acquires dialogue data from user questions and uses the multi-level intent recognition model to perform intent recognition on the dialogue data, obtaining recognition results for top-level intent, middle-level intent, and bottom-level intent.
[0152] If the above functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0153] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0154] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0155] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0156] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A multi-level customer intent recognition method for an intelligent customer service system based on phased learning, characterized in that, include: Acquire fully labeled data, weakly supervised data, and unlabeled data. Based on the fully labeled data and the weakly supervised data, set a loss function that emphasizes the top-level intent and establish a basic intent recognition model. Based on the basic intent recognition model, mid-level intent pseudo-labels and low-level intent pseudo-labels are generated for the weakly supervised data. Combined with dialogue coherence loss, the basic intent recognition model is trained using the unlabeled data to obtain a hierarchical intent recognition model. Based on the hierarchical intent recognition model, the underlying intent is optimized through a comparative learning task to obtain a multi-level intent recognition model. The dialogue data of the user's question is obtained, and the multi-level intent recognition model is used to perform intent recognition on the dialogue data to obtain the recognition results of the top-level intent, the middle-level intent and the bottom-level intent. Based on the fully labeled data and the weakly supervised data, a loss function emphasizing top-level intent is set to establish a basic intent recognition model. This includes setting a first loss function for the fully labeled data that includes top-level intent loss, mid-level intent loss, and bottom-level intent loss, wherein the weight of the top-level intent loss is greater than the weights of the mid-level intent loss and the bottom-level intent loss. A second loss function, incorporating the top-level intent loss, is applied to the weakly supervised data. A basic intent recognition model is established by jointly training the first loss function and the second loss function.
2. The method for multi-level customer intent recognition in an intelligent customer service system based on phased learning as described in claim 1, characterized in that: Generating mid-level and low-level intention pseudo-labels for the weakly supervised data based on the basic intention recognition model includes: using the basic intention recognition model to perform top-level intention prediction on the weakly supervised data, obtaining the top-level intention prediction confidence score, and filtering weakly supervised data whose top-level intention prediction confidence score is higher than a preset threshold to obtain the filtered weakly supervised data samples. Based on the real top-level intent labels of the selected weakly supervised data samples, random sampling is performed within the corresponding mid-level intent subclasses using a predefined hierarchical mapping relationship to generate mid-level intent pseudo-labels. Based on the mid-level intent pseudo-tags, random sampling is performed within the corresponding low-level intent subclass range to generate low-level intent pseudo-tags.
3. The method for multi-level customer intent recognition in an intelligent customer service system based on phased learning as described in claim 2, characterized in that: Training the basic intent recognition model using the unlabeled data, incorporating dialogue coherence loss, includes constructing a dialogue coherence prediction task and determining the intent coherence between adjacent dialogue statements. The unlabeled data is processed using the dialogue coherence prediction task to calculate the dialogue coherence loss. The mid-level intent pseudo-labels and the bottom-level intent pseudo-labels are combined with the dialogue coherence loss to train the basic intent recognition model, resulting in a hierarchical intent recognition model.
4. The method for multi-level customer intent recognition in an intelligent customer service system based on phased learning as described in claim 3, characterized in that: Based on the hierarchical intent recognition model, the optimization of the underlying intent through a comparative learning task includes clustering the feature representation output by the hierarchical intent recognition model to generate cluster pseudo-labels for the unlabeled data. Construct positive sample pairs, which include sample pairs with the same real labels and sample pairs with the same cluster pseudo labels; Construct negative sample pairs, which include sample pairs of different intent categories; The hierarchical intent recognition model is trained using the positive and negative sample pairs to optimize the classification boundary of the underlying intent.
5. The method for multi-level customer intent recognition in an intelligent customer service system based on phased learning as described in claim 4, characterized in that: Clustering the feature representation output by the hierarchical intent recognition model to generate cluster pseudo-labels for the unlabeled data includes extracting the feature vectors output by the feature sharing layer of the three classification heads in the hierarchical intent recognition model as clustering objects. The feature vectors are clustered using a clustering algorithm, with the number of clusters set to the total number of categories for the top-level intent, middle-level intent, and bottom-level intent. The clustering results are tracked over multiple consecutive training cycles, and pseudo-labels are assigned to samples whose clustering results are consistent and whose samples are close to the cluster center.
6. The method for multi-level customer intent recognition in an intelligent customer service system based on phased learning as described in claim 5, characterized in that: The construction of negative sample pairs includes constructing a first type of negative sample pair across the top-level intent category; Construct a second type of negative sample pair that belongs to the same top-level intent but to different mid-level intents; Based on the prediction confidence of the hierarchical intent recognition model, samples with high confidence but incorrect predictions are mined to construct a third type of negative sample pair; The first type of negative sample pairs, the second type of negative sample pairs, and the third type of negative sample pairs are combined to form hierarchical negative sample pairs.
7. A multi-level customer intent recognition system for intelligent customer service systems based on phased learning, employing the method described in any one of claims 1-6, characterized in that: The basic intent establishment module acquires fully labeled data, weakly supervised data, and unlabeled data. Based on the fully labeled data and the weakly supervised data, it sets a loss function that focuses on the top-level intent and establishes a basic intent recognition model. The hierarchical structure self-discovery module generates mid-level and low-level pseudo-labels for the weakly supervised data based on the basic intent recognition model, and trains the basic intent recognition model using the unlabeled data in conjunction with dialogue coherence loss to obtain a hierarchical intent recognition model. The fine-grained optimization module, based on the hierarchical intent recognition model, optimizes the underlying intent through a comparative learning task to obtain a multi-level intent recognition model; The intent recognition module acquires dialogue data from user questions and uses the multi-level intent recognition model to perform intent recognition on the dialogue data, obtaining recognition results for top-level intent, middle-level intent, and bottom-level intent. Based on the fully labeled data and the weakly supervised data, a loss function emphasizing top-level intent is set to establish a basic intent recognition model. This includes setting a first loss function for the fully labeled data that includes top-level intent loss, mid-level intent loss, and bottom-level intent loss, wherein the weight of the top-level intent loss is greater than the weights of the mid-level intent loss and the bottom-level intent loss. A second loss function, incorporating the top-level intent loss, is applied to the weakly supervised data. A basic intent recognition model is established by jointly training the first loss function and the second loss function.
8. A computer device, comprising: Memory and processor; The memory stores a computer program, characterized in that: when the processor executes the computer program, it implements the steps of the method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Multi-intention semantic recognition method in intelligent dialogue
CN115936011A
Potential user identification method, equipment, storage medium and device
CN116861314A