Dialogue Intention Classification Method, Device, and Non-Volatile Computer Storage Medium

By adjusting the feature vector modular length and joint training learning network, the problem of identifying unknown intention categories is solved, and the accurate classification of known intention categories and effective identification of unknown intention categories in the open intention classification task is realized, thereby reducing false positive errors.

CN114077666BActive Publication Date: 2025-07-29TOYOTA JIDOSHA KK +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010850324.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-21
Publication Date
2025-07-29
Estimated Expiration
2040-08-21

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively identify and classify dialogue data of unknown intention categories, especially in the open intention classification task where known intentions and unknown intentions coexist. Traditional methods are prone to introduce false positive errors and lack technical means to effectively utilize data generalization of known intention categories to unknown intention categories.

Method used

The first learning network is used to extract the first feature vector of the sentence, and by adjusting the modular length of the feature vector, combining the second learning network perceives the modular length information, the probability-related parameters of each known intention category are determined by combining the combination of the metric loss function and the classification loss function, and distinguishing between known intention and unknown intention through threshold comparison.

Benefits of technology

In the open intention classification task, unknown intention categories can be effectively identified, while ensuring the classification accuracy of known intention categories, reducing false positive errors, and improving the overall classification effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114077666B_ABST
    Figure CN114077666B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, apparatus, and non-volatile computer storage medium for classifying dialogue intents. The method includes receiving dialogue data and using a first learning network based on the data to extract a first feature vector of a sentence. Adjusting the norm of the first feature vector such that the smaller the minimum difference, the larger the norm, to obtain a second feature vector. Based on the second feature vector, using a second learning network to determine probability-related parameters for each known intent category, and jointly training the first learning network and the second learning network using a first loss function serving as a metric loss function and a second loss function representing classification loss. Comparing each probability-related parameter with a threshold to determine whether the dialogue belongs to an unknown intent category and which known intent category it belongs to. The method of the present disclosure can effectively detect unknown intent categories while ensuring the accuracy of classification of known intent categories when facing the task of open intent classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to methods and apparatuses for automatically processing and analyzing conversations, and more particularly, to methods and apparatuses for identifying and classifying conversation intents. Background Art

[0002] Artificial intelligence technologies, including dialogue robots, are widely used in the automatic processing and analysis of human conversations. Taking a dialogue robot as an example, it needs to effectively identify and classify conversation intents from conversation data. However, not all conversation intents are known, and there are also unknown conversation intents.

[0003] As shown in Table 1, dialogue robots often need to handle the task of open intent classification. That is to say, among the sentences included in the conversation data to be analyzed, some belong to known intent categories, and some cannot be identified and classified, and the intent label is an unknown intent. As an example of a known intent category, the intent label of "I want to take medicine" is "being sick", the intent label of "I have a headache" is "being sick", the intent label of "I'm hungry" is "eating", and the intent label of "I'm thirsty" is "drinking water". As an example of an unknown intent, from sentences such as "What's the weather like today?" and "Can you help me see what this is?", the intent category cannot be identified and classified by machine learning methods.

[0004] Table 1 Examples of Dialogue Open Intent Classification

[0005] The conversation content of the user Intention label I want to take medicine. Illness I have a slight headache. Illness I'm hungry. Eat I'm thirsty. Drink water …… …… What's the weather like today? Unknown intention Can you help me take a look at what this is? Unknown intention

[0006] Conversation data of unknown intent categories will interfere with the identification and classification of known intent categories, but currently there is a lack of effective methods for identifying and classifying unknown intent categories. In particular, in the face of the task of open intent classification, sentences of known intents and unknown intents coexist, and it is impossible to take into account the identification and classification of both known intents and unknown intents. Although traditional supervised classification methods have good classification effects on conversation data of known intent categories, due to the lack of prior knowledge of unknown intent categories, such traditional supervised classification methods will introduce false positive errors in the classification of conversation data of unknown intent categories, and there is currently a lack of technical means for effectively generalizing the conversation data of known intent categories to the identification and classification of unknown intent categories. Summary of the Invention

[0007] The present disclosure is provided to solve the above problems existing in the prior art.

[0008] There is a need for a conversation intent classification method, a conversation intent classification apparatus, and a non-volatile computer storage medium that can ensure the accuracy of the classification of known intent categories while effectively detecting unknown intent categories when facing the task of open intent classification.

[0009] According to a first aspect of the present disclosure, a method for classifying dialogue intents is provided. The method includes receiving dialogue data. The method further includes the following steps performed by a processor. Based on the received dialogue data, a first feature vector of a sentence can be extracted using a trained first learning network. Based on the minimum difference between the first feature vector and the representative feature vectors of each known intent category, the norm of the first feature vector can be adjusted such that the smaller the minimum difference, the larger the norm, to obtain a second feature vector. Based on the second feature vector, various probability-related parameters of each known intent category can be determined using a trained second learning network. Wherein, the second learning network is configured to perceive the norm of the second feature vector. The first learning network and the second learning network can be jointly trained using a first loss function and a second loss function representing classification loss. The first loss function can be defined such that the difference between the first feature vector and the representative feature vector of the known intent category to which it belongs is smaller than the representative value of the differences between the first feature vector and the representative feature vectors of other known intent categories. The determined various probability-related parameters of each known intent category are compared with a threshold. In the case where each probability-related parameter is less than the threshold, it is determined that the dialogue belongs to an unknown intent category; in the case where each probability-related parameter is above the threshold, it is determined that the dialogue belongs to the known intent category corresponding to the largest probability-related parameter.

[0010] According to a second aspect of the present disclosure, there is provided a dialogue intention classification device, which includes an interface and a processor. The interface may be configured to receive dialogue data. The processor may be configured to perform the following steps. Based on the received dialogue data, a first feature vector of a sentence may be extracted by using a trained first learning network. Based on the minimum difference between the first feature vector and the representative feature vectors of each known intention category, the norm of the first feature vector may be adjusted such that the smaller the minimum difference, the larger the norm, to obtain a second feature vector. Based on the second feature vector, each probability-related parameter of each known intention category may be determined by using a trained second learning network. Wherein, the second learning network is configured to sense the norm of the second feature vector. The first learning network and the second learning network may be jointly trained by using a first loss function and a second loss function representing classification loss. The first loss function may be defined such that the difference between the first feature vector and the representative feature vector of the known intention category to which it belongs is smaller than the representative value of the differences between the first feature vector and the representative feature vectors of other known intention categories. The determined probability-related parameters of each known intention category are compared with a threshold. In the case where each probability-related parameter is less than the threshold, it is determined that the dialogue belongs to an unknown intention category; in the case where each probability-related parameter is above the threshold, it is determined that the dialogue belongs to the known intention category corresponding to the maximum probability-related parameter.

[0011] According to a third aspect of the present disclosure, there is provided a non-volatile computer storage medium, on which executable instructions are stored, and when the executable instructions are executed by a processor, a dialogue intention classification method according to various embodiments of the present disclosure is implemented.

[0012] By using the dialogue intention classification method, the dialogue intention classification device, and the non-volatile computer storage medium according to various embodiments of the present disclosure, when facing the task of open intention classification, through the adjustment method of the norm of the first feature vector, the configuration of the second learning network to sense the norm, and the combined use of the metric loss function and the classification loss function for joint training, it is possible to effectively detect unknown intention categories while ensuring the accuracy of the classification of known intention categories. Description of the Drawings

[0013] In the drawings, which are not necessarily drawn to scale, the same reference numerals may describe similar components in different views. The same reference numerals with alphabetical suffixes or different alphabetical suffixes may represent different instances of similar components. The drawings generally illustrate various embodiments by way of example and not limitation, and are used together with the specification and the claims to explain the disclosed embodiments. Such embodiments are illustrative and not intended to be an exhaustive or exclusive embodiment of the device or method.

[0014] Figure 1(a) shows a flowchart of a dialogue intent classification method according to an embodiment of the present disclosure;

[0015] Figure 1(b) shows a flowchart of a learning network for constructing and training a dialogue intent classification method according to an embodiment of the present disclosure;

[0016] Figure 2 A framework diagram of a pre-trained language model for extracting a first feature vector of a sentence in a dialogue according to an embodiment of the present disclosure is shown;

[0017] Figure 3 A flowchart of a dialogue intent classification method according to another embodiment of the present disclosure is shown;

[0018] Figure 4 A configuration diagram of a dialogue intent classification system according to another embodiment of the present disclosure is shown; and

[0019] Figure 5 A block diagram of a dialogue intent classification device according to an embodiment of the present disclosure is shown. Detailed implementation manners

[0020] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the present disclosure will be described in detail below with reference to the accompanying drawings and specific implementation manners. The embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings and specific embodiments, but this is not a limitation to the present disclosure. The terms "first", "second", and "third" used in the present disclosure are only intended to distinguish the corresponding features, and do not represent such a need for sorting, nor necessarily represent only the singular form.

[0021] FIG. 1(a) shows a flowchart of a dialogue intention classification method according to an embodiment of the present disclosure. As shown in FIG. 1(a), the dialogue intention classification method starts from step 101: receiving dialogue data, where the dialogue contains multiple sentences. In step 102, a processor extracts a first feature vector of the sentence based on the received dialogue data by using a trained first learning network. The first feature vector contains semantic information of the sentence and can be extracted by a first learning network with various configurations. For example, words can be represented by discrete spatial vectors, such as the encoding in the bag-of-words model. Another example is that the first learning network can convert discrete vectors in a high-dimensional space into a dense vector distributed representation in a low-dimensional space, which can relate the distance relationship between words and the semantic similarity of word meanings, and each dimension can contain specific meanings, thereby containing more semantic information and suppressing the increase in the dimension of the spatial vector and the increase in spatial overhead compared with the discrete spatial vector representation. In some embodiments, a language model with dynamic word vectors can be adopted, so as to dynamically adjust the feature vector of the word according to different contexts. Thus, in cases such as polysemy that cannot be solved by static word vectors (each word has a fixed word vector), the feature vectors of the same word in different contexts can contain rich semantics and accurately reflect the influence of different contexts.

[0022] In step 103, the processor adjusts the norm of the first feature vector based on the minimum difference between the first feature vector and the representative feature vectors of each known intention category, such that the smaller the minimum difference, the larger the norm, to obtain a second feature vector. Although the first feature vector can contain richer semantic information through a language model with dynamic word vectors, the inventors found that it still cannot effectively distinguish unknown intention categories from known intention categories alone. By introducing difference information (the difference information can be implemented as various defined distance information, difference information, correlation information, etc.) on the basis of the first feature vector and guiding the adjustment of the norm of the feature vector with the difference information, the norm of the feature vector can introduce and take into account the difference information in the form of an adjustment coefficient. The adjustment coefficient can represent the degree of proximity of each data to the known intention. The larger the adjustment coefficient, the closer it is to the known intention category, and the smaller it is, the closer it is to the unknown intention category. In this way, the second feature vector (hereinafter its example is also referred to as "meta-embedding") after the norm adjustment can learn the data relationship of being close within the class and far between classes (especially far between the known intention category and the unknown intention category) at a deep level.

[0023] In step 104, the processor determines respective probability-related parameters for each known intent category based on the second feature vector by using a second learning network configured to be able to sense the norm information of the feature vector. The second feature vector whose norm is adjusted based on the difference information in step 103, in cooperation with the second learning network configured to be able to sense (capture) the norm information of the feature vector, can obtain probability-related parameters (such as classification confidence) that are easy to distinguish.

[0024] Then in step 105, the processor compares the determined respective probability-related parameters for each known intent category with a threshold. In the case where all the probability-related parameters are less than the threshold, it is determined that the conversation belongs to an unknown intent category (step 106); in the case where all the probability-related parameters are above the threshold, it is determined that the conversation belongs to the known intent category corresponding to the maximum probability-related parameter (step 107).

[0025] FIG. 1(b) shows a flowchart of a training method of a learning network adopted by the dialogue intent classification method according to an embodiment of the present disclosure. The learning network includes a first learning network and a second learning network.

[0026] The training method may start from step 108: constructing a first learning network and a second learning network, where the first learning network may be configured to extract a first feature vector of a sentence, and the second learning network may be configured to sense the norm information of the second feature vector. As described above, the first learning network may adopt various network structures. In some embodiments, based on the received dialogue data, a pre-trained language model may be used to extract respective third feature vectors of each language symbol considering the context of each language symbol; and based on the extracted respective third feature vectors, the first feature vector of the sentence may be determined through comprehensive operations. In some embodiments, the pre-trained language model includes but is not limited to the BERT language model, the RoBERTa language model, various recurrent neural networks (RNNs, such as but not limited to long short-term memory neural networks (LSTMs)), XLNet, etc. The pre-trained language model may be pre-trained through multiple tasks to capture different context information of words. After pre-training, some parameters in the first learning network may be fixed, and in the subsequent training step 111 together with the second learning network, only the remaining parameters in the first learning network are determined and adjusted (fine-tuned). In this way, the computational amount during the training process can be significantly reduced while taking into account the training effect.

[0027] In step 109, training corpora of known intent categories are received. For the dialogue data of unknown intent categories, there is a lack of training corpora. Through the cooperation of the definition of the second feature vector and the configuration of the second learning network, as well as the special definition of the first loss function, in this training method, the training corpora of known intent categories can be used to jointly train the first learning network and the second learning network. The trained second learning network can effectively distinguish unknown intent categories and accurately classify known intent categories in the application scenario of dialogue open intent classification.

[0028] In step 110, each training corpus can be loaded as a training sample. In step 111, at least some parameters of the first learning network and the second learning network can be determined based on the training sample. Among them, the first learning network using the pre-trained language model can directly use some pre-trained network layers, and only the parameters of the remaining network layers are determined and adjusted in step 111. In step 112, the first learning network and the second learning network can be jointly trained for the first loss function and the second loss function to verify and adjust the network parameters. Among them, the second loss function can adopt various forms of loss functions representing classification loss, including but not limited to the square loss function, the cross-entropy loss function, etc. The first loss function is defined such that the difference between the first feature vector and the representative feature vector of its known intent category is smaller than the representative value of each difference from the representative feature vectors of other intent categories. For example, a loss boundary can be set (such as in the form of a set value, a boundary penalty term, etc.). In this way, through the constraint of the first loss function, the learning network learns features that are close within the class and far between classes, which the second learning network must consider when learning the decision boundary.

[0029] Note that although step 111 and step 112 are shown one after the other in FIG. 1(b), their execution order is not limited to this, as long as it is considered that both the first loss function and the second loss function adjust at least some parameters of the first learning network and the second learning network.

[0030] In step 113, it is determined whether all training samples have been processed. If so, in step 114, the trained first learning network and the second learning network are output; if not, it returns to step 110 to continue loading the next training sample for training. In some embodiments, a batch training method can be adopted, such as a mini-batch training method. Correspondingly, according to the specific training method adopted, steps 110, 111, 112, and 113 will be adjusted accordingly, which will not be elaborated here.

[0031] Figure 2 A framework diagram of a pre-trained language model for extracting the first feature vector of a sentence in a dialogue according to an embodiment of the present disclosure is shown. As Figure 2As shown, for a sentence 201 input by a user, S = {word1, word2, …, wordn}, where n is the total number of words in the sentence. Here, words are used as examples of language symbols, and it can also be generalized to various language symbols (such as words, etc.) to implement this pre-trained language model. In the following text, in combination with Figure 2 BERT pre-trained language model is used as an example of the pre-trained language model for illustration. However, it should be noted that the pre-trained language model is not limited to this, and various implementation methods such as RoBERTa language model, various recurrent neural networks (RNN, such as but not limited to long short-term memory neural network (LSTM)), XLNet, etc. can also be used. The description of the sentence feature vector extraction of these pre-trained language models can also be applied through adaptive modification in combination with the description of the BERT pre-trained language model in the following text, which will not be elaborated here.

[0032] S can be encoded in the way of BERT, and then pass through the transformation layer of BERT (also called the hidden encoding layer), the first transformation layer 202, ……, the twelfth transformation layer 203, to obtain the respective word feature vectors of each word [CLS, T1, T2, … T n , where CLS is used as an element and serves as a classification flag, and T i represents the feature vector of the i-th word. Figure 2 Twelve transformation layers are shown as an example, but the number of transformation layers is not limited to this. The BERT pre-trained language model can include 12 - 24 transformation layers.

[0033] The pooling layer 204 can be used to perform a pooling operation on the extracted respective word feature vectors (also called the third feature vectors in this disclosure), such as but not limited to an average pooling operation, to determine the fourth feature vector x of the sentence, x = average pooling([CLS, T1, T2, …, T n ) ∈ R H , where H is the number of hidden layer neurons. In Figure 2 the pooling layer 204 is used as an example, and based on the extracted respective third feature vectors, the first feature vector of the sentence can also be determined through other forms of comprehensive (composite) operations. This comprehensive operation is configured to integrate the third feature vectors of each word to obtain a sentence-level feature vector representation. In some embodiments, the determined fourth feature vector of the sentence can also be used as the first feature vector of the sentence. In some embodiments, as Figure 2 shown, a dense layer 205 can be set downstream of the pooling layer 204, so as to process the fourth feature vector of the sentence using the dense layer 205 to obtain the first feature vector z of the sentence (also identified as “BERT embedding” in this disclosure), z = f θ (x) ∈ R D, where θ is a parameter of the dense layer 205 and D is the dimension of the feature vector.

[0034] In some embodiments, the BERT language model can be pre-trained through two tasks: masked language prediction and next sentence prediction, so that it can capture different context information of words, enabling the feature vectors of words to be dynamically adjusted according to different contexts, thereby containing richer semantic information and external knowledge, and enhancing the representation effect of features. By introducing the dense layer 205, the discrete vector of the sentence in the higher-dimensional space output by the pooling layer 205 can be converted into a dense feature vector of the sentence in the lower-dimensional space. This dense feature vector of the sentence in the lower-dimensional space has the characteristics of distributed representation, can relate the difference (distance) relationship between words and the semantic similarity of word meanings, and each dimension can contain specific meanings. Therefore, it can contain richer information, significantly suppress the increase in space overhead, improve computational efficiency, and at the same time improve the prediction ability of the model.

[0035] In some embodiments, the BERT language model may include 12 - 24 transformation layers, and only the last 1 transformation layer and the dense layer are trained during the joint training of the first learning network and the second learning network. That is to say, the other transformation layers can directly use the pre-trained parameters without further training. Taking the BERT language model with 12 transformation layers as an example, during the subsequent joint training, actually only the last transformation layer (i.e., the 12th transformation layer 203) and the dense layer 205 need to be trained. In this way, the number of training parameters can be significantly reduced, the computational load of training can be reduced, and the training speed can be increased. The BERT language model can extract sentence-level feature vectors for the application scenario of dialogue open intent classification through fine-tuning downstream.

[0036] Figure 3 The flowchart of a dialogue intent classification method according to another embodiment of the present disclosure is shown. Figure 3 Taking the pre-trained BERT language model as an example of the pre-trained language model, the cluster center vector of the known intent category as an example of the representative feature vector of the known intent category, correspondingly, taking the Euclidean distance between the first feature vector and the cluster center vector of the known intent category as an example of the difference between the first feature vector and the representative feature vector, and taking the cosine classifier as an example of the second learning network, the process of the dialogue intent classification method according to the embodiments of the present disclosure is described. It should be noted that other pre-trained language models can be used, the representative feature vector of the known intent category can also take other forms, such as the cluster median vector, etc., the difference can also be implemented as correlation, difference, etc., the Euclidean distance can also be modified to the Bhattacharyya distance, etc., and the second learning network can also adopt other configurations, which will not be elaborated here.

[0037] As Figure 3 shown, a statement input by a user is fed into a pre-trained BERT language model 301 to obtain a BERT embedding 302, that is, the first feature vector z = f θ (x) ∈ R D .

[0038] First, the cluster center vectors of each known intent category can be calculated according to formula (3.1) as the representative feature vectors of the known intent category, where the cluster center vectors are obtained by averaging the feature vectors of all the data of the known intent category:

[0039]

[0040] where c k represents the cluster center vector of the k-th category, S k represents the data set with label k, x represents the statement sample, y represents the label, and |·| is used to count the size of the set.

[0041] Next, the distance from each data point (i.e., each first feature vector z) to the nearest cluster center can be calculated according to formula (3.2):

[0042]

[0043] where d min represents the minimum distance from the data point z to all the cluster centers , ||·||2 represents calculating the Euclidean distance, represents the minimum value of all K distances.

[0044] The Euclidean distance information can be introduced based on the first feature vector z, and the norm adjustment of the first feature vector z can be guided by the Euclidean distance information via an adjustment coefficient. The adjustment coefficient can represent the degree of proximity of each data to the known intent. The larger the adjustment coefficient, the closer it is to the known intent category, and the smaller it is, the closer it is to the unknown intent category. Through the adjustment coefficient, the feature vector can be guided to learn the data relationship of being close within the class and far between classes. For example, the reciprocal of the minimum distance d min obtained in the previous step can be set as the adjustment coefficient and added to the original first feature

[0045] vector z to obtain the second feature vector z meta ( Figure 3 labeled as "meta-embedding 303" in):

[0046] z meta = (1 / d min )·z Formula (3.3)

[0047] Among them, 1 / d min can represent the degree of proximity of each data point to a known intention category. The larger the adjustment coefficient, the closer it is to the known intention category, and the smaller it is, the closer it is to the unknown intention category. By determining the magnitude characteristic of the second feature vector z meta through the adjustment coefficient, this characteristic is conducive to learning the deep data relationship of being close within the class (close within the known intention category) and far between classes (especially far between the known intention category and the unknown intention category), enabling the subsequent second learning network ( Figure 3 illustrated as the cosine classifier 304 in) to obtain more distinguishable probability-related parameters ( Figure 3 shown as the classification probability 306 as an example in).

[0048] For example Figure 3 as shown, based on the second feature vector z meta incorporating distance information via the adjustment coefficient, the cosine classifier 304 is used to convert its magnitude (including distance information) into classification probability 306 information.

[0049] The cosine classifier 304 is essentially a layer of weights of a neural network. This weight can learn the magnitude information of the second feature vector z meta through network training. According to formula (3.4), by calculating the cosine similarity between the second feature vector z meta and the classification weight, the classification score can be obtained:

[0050]

[0051] where, sim k represents the cosine similarity between the second feature vector z meta and the classification weight vector of the k-th class . τ is a scalar value that can be learned, and the cosine similarity can be calculated using the dot product.

[0052] In some embodiments, as shown in the right-hand expression of formula (3.4), the second feature vector z meta and the classification weight vector of the k-th class can be normalized, such as but not limited to L1 norm normalization, L2 norm normalization, etc.

[0053] The following takes L2 norm normalization as an example of the normalization process for illustration. and both represent the vectors of the second feature vector z meta and the classification weight vector of the k-th class after L2 norm normalization. Specifically, and are calculated as follows:

[0054]

[0055]

[0056] Among them, ||·|| represents the norm of the feature, and the second feature vector z can be normalized by the L2 norm using a non-linear squeezing function meta The non-linear squeezing function can compress a vector with a larger norm to a length slightly less than 1 and compress a vector with a smaller norm to a length close to 0. The features obtained through such processing can transform the feature vectors of unknown intention categories with smaller norms into feature vectors with extremely small lengths, thereby obtaining very low classification scores that are easy to distinguish by a threshold. In some embodiments, as defined by formula (3.6), the classification weight vector can also be normalized by the L2 norm, so as to eliminate the influence of the norm of the classification weight vector on the classification result and make the classification more accurate.

[0057] In some embodiments, the softmax function can be applied to the classification scores obtained using the cosine classifier 304 to obtain classification probabilities 306, and then based on a confidence threshold, classification of known intention categories and unknown intention categories is performed, thereby completing the classification prediction 305, as Figure 3 shown.

[0058] The first loss function and the second loss function can be used to jointly train some network layers of the pre-trained BERT language model 301 and the cosine classifier 304. In some embodiments, according to formula (3.7), the first loss function Loss d and the second loss function Loss ce can be combined as the total loss function Loss for joint training. Among them, the first loss function Loss d can be used as a metric learning loss function, while the second loss function Loss ce can be used as a classification loss function. By combining the classification loss function with the metric learning loss function for joint training, while ensuring the classification performance, a feature representation that can perceive distance information can be learned.

[0059] Loss = Loss ce + λ·Loss d Formula (3.7)

[0060] Among them, λ is a scalar value used to balance the two loss functions.

[0061] In some embodiments, the second loss function Loss ceAs the classification loss function, the cross-entropy loss function defined by the following formulas (3.8) and (3.9) can be adopted:

[0062]

[0063]

[0064] Wherein, represents the linear output layer, i represents the sample serial number, N represents the total number of samples, K represents the number of known intention categories, y i represents the known intention category label of the i-th sample and y i ∈{1, 2, …, K}, m1 represents the set loss boundary, j represents the serial number of the known intention category, represents the weight parameter of the known intention category identified by the label y i and represents the weight vector of the j-th category, represents the bias of the known intention category identified by the label y i and b j represents the bias of the j-th category, represents the second feature vector (also referred to as "meta-feature") extracted from the i-th sample. This meta-feature can be calculated by using the softmax function to obtain the classification probability 306, and the cross-entropy loss is defined as maximizing the probability that the meta-feature is assigned to the intention category it belongs to.

[0065] In some embodiments, the first loss function Loss d can be used as a metric learning loss function, which can be defined by various formulas as long as it realizes the following physical meaning: the difference (such as Euclidean distance) between the first feature vector and the representative feature vector of the known intention category it belongs to is smaller than the representative value (such as mean value) of the differences (such as Euclidean distance) between it and the representative feature vectors of other intention categories, for example, smaller than at least a certain loss boundary threshold. In this way, by introducing the constraint of this metric learning loss function Loss d into the total loss function, the first learning network (correspondingly, the first feature vector) can learn the deep data relationship of being close within the class and far between classes. Since the first loss function Loss d provides the constraint that the difference between the first feature vector and the known intention category and the unknown intention category is at least the loss boundary threshold, it is relatively easier to determine whether it belongs to the known intention category or the unknown intention category by the distance between the data and the known intention category.

[0066] In some embodiments, the first loss function Loss d can be defined by the following formula (1):

[0067]

[0068] Among them, Loss1 represents the first loss function Loss d or its first component, ||·||2 represents the Euclidean distance, i represents the sample serial number, N represents the total number of samples, K represents the number of known intention categories, k represents the known intention category serial number, y i represents the known intention category label of the i-th sample and y i ∈{1, 2, …, K}, m1 represents the set loss boundary, z i represents the first feature vector extracted for the i-th sample, c k represents the cluster center vector of the k-th known intention category. Formula (1) provides the following physical meaning: the Euclidean distance from the first feature vector to the cluster center vector of its known intention category is at least smaller than the mean of the distances to the cluster center vectors of other intention categories by m1. Through the constraint of the first loss function Loss d , the first feature vector has the characteristics of being close within the class and far between classes. Since the Euclidean distance from the first feature vector to the known intention category and the unknown intention category differs by at least the boundary value m1, it is relatively easier to determine whether it belongs to the known class or the unknown class by the distance of the data to the known intention category.

[0069] In some embodiments, the first loss function Loss d can also be defined by using the combination of the components Loss1, Loss2, and Loss3 of the first loss function defined by the formula (1), the following formula (2), and formula (3):

[0070]

[0071] Among them, Loss2 represents the second component of the first loss function, s1 is a scaling factor, is the angle between the first feature vector z i and the weight vector i of the known intention category identified by the label y , j represents the known intention category serial number, θ j is the angle between the first feature vector z i and the weight vector W j of the j-th class, and m2 is the cosine distance margin constant;

[0072]

[0073] Among them, Loss3 represents the third component of the first loss function, s2 is a scaling factor, m3 is the angular distance margin constant, and other parameters the same as those in formula (2) have the same definition and will not be elaborated here.

[0074] By combining Loss1, Loss2, and Loss3, decision boundary constraints can be imposed on the first feature vector in terms of the three metrics of Euclidean distance, cosine distance, and angular distance, further promoting proximity within classes and separation between classes, enabling more accurate classification of known intent classes and effective detection of unknown intent classes when facing the application scenario of open intent classification in conversations. This beneficial effect has been verified in comparative experiments using three publicly available standard datasets, which will be described in detail later and will not be elaborated here.

[0075] Figure 4 A configuration diagram of a dialogue intent classification system 400 according to another embodiment of the present disclosure is shown. As Figure 4 shown, the dialogue intent classification system 400 can at least include an intent classification unit 406 configured to: receive a trained first learning network and a second learning network as well as dialogue data; based on the received dialogue data, use the trained first and second learning networks to determine respective probability-related parameters of the dialogue data with respect to each known intent class; and based on the respective probability-related parameters, obtain an intent classification result of the dialogue, which can be open-ended, such as whether the dialogue belongs to a known intent class or an unknown intent class, and which known intent class it belongs to, etc. The first and second learning networks can be trained elsewhere and fed into the intent classification unit 406. In some embodiments, the first learning network can be configured to extract a first feature vector of a sentence based on the dialogue data, and the second learning network can be configured to determine respective probability-related parameters of each known intent class based on a second feature vector obtained after adjusting the norm of the first feature vector. Accordingly, the intent classification unit 406 can be configured to adjust the norm of the first feature vector based on the minimum difference between the first feature vector and the representative feature vectors of each known intent class, such that the smaller the minimum difference, the larger the norm, and thus the norm of the obtained second feature vector can represent the proximity of the corresponding dialogue data to the known intent classes, thereby enabling the second feature vector to incorporate proximity information. This characteristic is beneficial for learning deep data relationships of proximity within classes (proximity within known intent classes) and separation between classes (especially separation between known intent classes and unknown intent classes), enabling the subsequent second learning network to obtain more distinguishable probability-related parameters. In some embodiments, the intent classification unit 406 can be configured to: compare the determined respective probability-related parameters of each known intent class with a threshold, and if all the probability-related parameters are less than the threshold, determine that the dialogue belongs to an unknown intent class, and if all the probability-related parameters are above the threshold, determine that the dialogue belongs to the known intent class corresponding to the maximum probability-related parameter.

[0076] In some embodiments, the dialogue intention classification system 400 may further include a learning network construction unit 401 configured to construct a first learning network and a second learning network. For example, the first learning network may be constructed based on a pre-trained language model. The construction of the first and second learning networks has been described in detail in other embodiments of the present disclosure and will not be elaborated here. Pre-training samples can be obtained from the pre-training sample database 403, and the pre-training unit 402 is used to complete the pre-training of the first learning network. The pre-trained first learning network and the constructed second learning network can be transmitted to the training unit 404 to complete the training of the first and second learning networks using the training samples obtained from the training sample database 405. After the pre-training is completed, some parameters in the first learning network can be maintained in the state after the pre-training is completed, and in the subsequent training steps together with the second learning network, only the remaining parameters in the first learning network and the parameters of the second learning network are determined and adjusted (fine-tuned). In this way, the computational amount during the training process can be significantly reduced while taking into account the training effect.

[0077] Each of the above units can be respectively configured to execute the corresponding pre-training process, training process, intention classification process, etc. described in other embodiments of the present disclosure, which will not be elaborated here.

[0078] Figure 5 The block diagram of the dialogue intention classification device 500 according to an embodiment of the present disclosure is shown. As Figure 5 shown, the network training device 501 is configured as a device independent of the dialogue intention classification device 500. The former is configured to construct and train the first and second learning networks. The trained first and second learning networks can be fed to the dialogue intention classification device 500 via the communication interface 503 for its use. However, this is only an example, and these two devices can also be integrated into the same device.

[0079] The dialogue data acquisition device 502 can be configured to acquire dialogue data. For example, it can include a microphone, an analog-to-digital converter, a filter, etc. The dialogue data acquired by it can be transmitted to the dialogue intention classification device 500 via the communication interface 503. In some embodiments, the dialogue data acquisition device 502 can be integrated with the dialogue intention classification device 500 in the same device, such as but not limited to a smart wearable device, a smart service robot, a smart phone, etc.

[0080] The dialogue intention classification device 500 can be a dedicated computer or a general-purpose computer. For example, the dialogue intention classification device 500 can be a customized computer to execute the dialogue data acquisition and dialogue data processing tasks. As Figure 5 shown, the dialogue intention classification device 500 can include a communication interface 503, a processor 504, a memory 505, a storage 506, and a display 507.

[0081] The communication interface 503 may include a network adapter, a cable connector, a serial connector, a USB connector, a parallel connector, a high-speed data transfer adapter (such as fiber optic, USB 3.0, Thunderbolt interface, etc.), a wireless network adapter (such as a WiFi adapter), a telecommunications (3G, 4G / LTE, 5G, etc.) adapter, and the like. The dialogue intention classification device 500 may be connected to other components through the communication interface 503, such as but not limited to the network training device 501, the dialogue data acquisition device 502, and the like. In some embodiments, the communication interface 503 receives dialogue data from the dialogue data acquisition device 502. In some embodiments, the communication interface 503 may also receive, from the network training device 501, such as the trained first learning network and the second learning network.

[0082] The processor 504 may be a processing device including more than one general-purpose processing device, such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), and the like. More specifically, the processor may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor running other instruction sets, or a processor running a combination of instruction sets. The processor may also be more than one dedicated processing device, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a system on a chip (SoC), and the like. The processor 504 may be communicatively coupled to the memory 505 and be configured to execute computer-executable instructions stored thereon to perform a dialogue intention classification processing procedure such as described in various embodiments of the present disclosure.

[0083] The memory 505 / storage 506 may be a non-transitory computer-readable medium, such as a read-only memory (ROM), a random access memory (RAM), a phase change random access memory (PRAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), an electrically erasable programmable read-only memory (EEPROM), other types of random access memory (RAM), a flash drive or other forms of flash memory, a cache, a register, a static memory, a compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), or other optical memory, a cassette tape or other magnetic storage device, or any other possible non-transitory medium used to store information or instructions accessible by a computer device, and the like.

[0084] In some embodiments, the memory 506 may store the trained first and second learning networks and data (such as original conversation data, first feature vectors, intermediate parameters for adjusting the norm of the first feature vectors, second feature vectors, various probability-related parameters for each known intent category, comparison thresholds for each probability-related parameter, etc.), data received, used, or generated while executing the computer program, and the like. In some embodiments, the memory 505 may store computer-executable instructions, such as more than one conversation intent classification program.

[0085] In some embodiments, the processor 504 may be further configured to cause a corresponding output device to provide a corresponding service based on the intent classification result of the conversation. For example, the output device may be Figure 5 the display 507 shown in, but not limited to, a speaker, a mobile driving device, etc. For example, when the intent classification result of the conversation is "eating", the processor 504 may be configured to cause the display 507 to display a list of nearby restaurants for the user to select, and may also cause the display 507 to display a navigation map of the selected restaurant and play a navigation voice signal via the speaker. Again, for example, when the intent classification result of the conversation is "being ill", the processor 504 may be configured to open an online medical consultation platform, display the display interface of the consultation platform on the display 507, and cause the speaker (not shown) to play the voice of the consultation, such as "Please describe the location of your pain", "Do you have a designated doctor?", "Please upload the inspection report", etc.

[0086] In some embodiments, the display 507 may include a liquid crystal display (LCD), a light-emitting diode display (LED), a plasma display, or any other type of display, and provide a graphical user interface (GUI) presented on the display for user input and image / data display. The display may include many different types of materials (such as plastic or glass) and may be touch-sensitive to receive commands from the user. For example, the display may include a substantially rigid touch-sensitive material (such as Gorilla glassTM) or a substantially flexible (such as Willow glassTM) touch-sensitive material.

[0087] According to the present disclosure, the network training device 501 may have the same or similar structure as the conversation intent classification device 500. In some embodiments, the network training device 501 includes a processor and other components configured to train the first and second learning networks.

[0088] Comparative experiments were conducted on the general model and the dialogue intention classification method according to various embodiments of the present disclosure using three publicly available standard datasets, namely the StackOverflow technical question dataset, the FewRel relation extraction dataset, and the OOS dialogue dataset provided by Clinc Corporation. Among them, StackOverflow contains 20 technical question intention labels, FewRel contains 80 relation intentions, and the OOS dataset covers 150 intentions in 20 domains and 1,200 out-of-domain data.

[0089] The macro F1 value was used as the evaluation metric for each model, and each dataset was divided into independent training sets, validation sets, and test sets. The training set and the validation set only contain known class intentions, and the test set contains known class intentions and open intentions. The validation set was used to screen the optimal parameters. The general model adopted a traditional supervised classification model and a cross-entropy loss function. The dialogue intention classification method according to various embodiments of the present disclosure includes five variants of classification methods, called Dialogue Intention Classification Method 1, Dialogue Intention Classification Method 2, Dialogue Intention Classification Method 3, Dialogue Intention Classification Method 4, and Dialogue Intention Classification Method 5. Specifically, Dialogue Intention Classification Method 1 utilized the cross-entropy loss function based on the classification model shown in Figure 3 , and Dialogue Intention Classification Methods 2-5 were jointly trained using the cross-entropy loss function and the metric loss functions according to various embodiments of the present disclosure based on the classification model shown in Figure 3 . Among them, the metric loss function adopted by Dialogue Intention Classification Method 2 is the loss function shown in Formula (1), the metric loss function adopted by Dialogue Intention Classification Method 3 is the loss function shown in Formula (2), the metric loss function adopted by Dialogue Intention Classification Method 4 is the loss function shown in Formula (3), and the metric loss function adopted by Dialogue Intention Classification Method 5 is the combination of the loss functions of each metric shown in Formulas (1)-(3).

[0090] Table 1 shows the results of the comparative experiment, where Dialogue Intention Classification Methods 1-5 are abbreviated as Example Methods 1-5.

[0091] Table 1 Comparative Experiment Results of Dialogue Open Intention Classification Methods

[0092]

[0093] The comparison experiment results show that, compared with the ordinary model, the open intention classification method of each embodiment of the present disclosure can significantly improve the classification performance under different proportions of known class intentions, has good robustness, and maintains stable performance in datasets of different scenarios. In particular, by jointly training with the cross-entropy loss function and the metric loss function according to various embodiments of the present disclosure, the classification performance and robustness of the same classification model can be significantly improved; furthermore, by using the combination of several independent metric loss functions as the metric loss function for joint training with the cross-entropy loss function, compared with using a single metric loss function for joint training with the cross-entropy loss function, the classification performance and robustness of the classification model can be significantly improved.

[0094] In addition, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present disclosure having equivalent elements, modifications, omissions, combinations (e.g., schemes that cross various embodiments), adaptations or changes. The elements in the claims will be broadly interpreted based on the language employed in the claims and are not limited to the examples described in this specification or during the implementation of this application, and the examples will be construed as non-exclusive. Thus, this specification and the examples are intended to be considered only as examples, and the true scope and spirit are indicated by the full scope of the following claims and their equivalents.

[0095] The order of the various steps in the present disclosure is merely exemplary and not restrictive. Without affecting the implementation of the present disclosure (without disrupting the logical relationship between the required steps), the execution order of the steps can be adjusted, and the various embodiments obtained after the adjustment still fall within the scope of the present disclosure.

[0096] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of them) can be used in combination with each other. For example, those of ordinary skill in the art can use other embodiments when reading the above description. Additionally, in the above detailed description, the various features can be grouped together to simplify the present disclosure. This should not be construed as an intention that the disclosed features not claimed are necessary for any claim. On the contrary, the subject matter of the present invention may be less than all the features of a particular disclosed embodiment. Thus, the following claims are incorporated herein as examples or embodiments into the detailed description, where each claim independently serves as a separate embodiment, and considering these embodiments, they can be combined with each other in various combinations or permutations. The scope of the present invention should be determined with reference to the appended claims and the full scope of the equivalents to which these claims are entitled.

Claims

1. A method for classifying dialogue intents, characterized in that, Including: Receiving data of a conversation; Using a trained first learning network by a processor to extract a first feature vector of a sentence based on the received conversation data; Adjusting a norm of the first feature vector by the processor based on a minimum difference between the first feature vector and representative feature vectors of respective known intent categories, such that the smaller the minimum difference, the larger the norm, to obtain a second feature vector; Determining respective probability-related parameters of respective known intent categories by the processor based on the second feature vector using a trained second learning network, the second learning network being configured to sense the norm of the second feature vector, wherein the first learning network and the second learning network are jointly trained using a first loss function and a second loss function representing a classification loss, the first loss function being defined such that a difference between the first feature vector and a representative feature vector of its belonging known intent category is smaller than a representative value of respective differences between the first feature vector and representative feature vectors of other known intent categories, and the probability-related parameters being parameters related to a conversation intent classification probability; Comparing, by the processor, the determined respective probability-related parameters of respective known intent categories with a threshold, and determining that the conversation belongs to an unknown intent category if all the probability-related parameters are less than the threshold, and determining that the conversation belongs to the known intent category corresponding to the maximum probability-related parameter if all the probability-related parameters are above the threshold.

2. The method for classifying dialogue intentions according to claim 1, wherein Using a trained first learning network to extract a first feature vector of a sentence based on the received conversation data further includes: Using a pre-trained language model to extract respective third feature vectors of respective language tokens considering contexts of the respective language tokens based on the received conversation data; Determining the first feature vector of the sentence through a comprehensive operation based on the extracted respective third feature vectors.

3. The method for classifying dialogue intentions according to claim 2, characterized in that, Determining the first feature vector of the sentence through a comprehensive operation based on the extracted respective third feature vectors further includes: Performing a pooling operation on the extracted respective third feature vectors to determine a fourth feature vector of the sentence; and Processing based on the fourth feature vector of the sentence using a dense network layer to obtain the first feature vector of the sentence.

4. The method for classifying dialogue intentions according to claim 3, characterized in that, The pre-trained language model includes a BERT language model, the BERT language model includes 12 - 24 transformation layers, and only its last 1 transformation layer and the dense network layer are trained during the joint training of the first learning network and the second learning network.

5. The method for classifying conversation intents according to claim 1, characterized in that The representative feature vectors of known intent categories are cluster center vectors of the known intent categories, the difference between the first feature vector and the representative feature vector includes a distance between the first feature vector and the representative feature vector, and the second learning network includes a cosine classifier.

6. The method for classifying dialogue intentions according to claim 5, characterized in that, Based on the second feature vector, using the trained second learning network to determine the various probability-related parameters of each known intent category further includes: normalizing the second feature vector and normalizing the weight of the cosine classifier; based on the normalized second feature vector, using the cosine classifier after weight normalization to determine the various probability-related parameters of each known intent category.

7. The method for classifying dialogue intentions according to claim 5, characterized in that, The second loss function includes a cross entropy loss function, and the first loss function is defined using the following formula (1): Among them, Loss1 represents the first loss function or its first component, ||·||2 represents the Euclidean distance, i represents the sample serial number, N represents the total number of samples, K represents the number of known intention categories, k represents the known intention category serial number, and y i represents the known intention category label of the i-th sample and y i ∈{1,2,…,K}, m1 represents the set loss boundary, z i represents the first feature vector extracted for the i-th sample, and c k represents the cluster center vector of the k-th known intention category.

8. The method for classifying dialogue intentions according to claim 7, characterized in that, The first loss function is defined by combining the components Loss1, Loss2, and Loss3 of the first loss function defined by formula (1), the following formula (2), and formula (3): Among them, Loss2 represents the second component of the first loss function, s1 is the scaling factor, is the first eigenvector z i with label y i Weight vectors for the identified known intent categories The angle between them, j represents the known intention category number, θ j is the first eigenvector z i and the weight vector W of the jth class j The angle between them, m2 is the cosine distance marginal constant; Among them, Loss3 represents the third component of the first loss function, s2 is the scaling factor, and m3 is the angular distance marginal constant.

9. A dialogue intention classification device, characterized in that, include: An interface configured to receive data for a conversation; as well as Processor, configured as: Based on the received conversation data, using the trained first learning network, extracting a first feature vector of the sentence; Based on a minimum difference between the first feature vector and representative feature vectors of each known intent category, adjusting the modulus of the first feature vector so that the smaller the minimum difference, the larger the modulus, to obtain a second feature vector; Based on the second feature vector, determining probability-related parameters for each known intent category using a trained second learning network, where the second learning network is configured to perceive the modulus of the second feature vector, wherein the first learning network and the second learning network are jointly trained using a first loss function and a second loss function representing classification loss, wherein the first loss function is defined such that a difference between the first feature vector and a representative feature vector of the known intent category to which it belongs is smaller than representative values of differences between the first feature vector and representative feature vectors of other known intent categories, and the probability-related parameters are parameters related to the probability of dialog intent classification; Compare the probability-related parameters of each known intent category determined with the threshold. When all probability-related parameters are less than the threshold, it is determined that the conversation belongs to an unknown intent category. When all probability-related parameters are above the threshold, it is determined that the conversation belongs to the known intent category corresponding to the maximum probability-related parameter.

10. The dialogue intention classification device according to claim 9, characterized in that, Extracting a first feature vector of a sentence based on the received conversation data using the trained first learning network further includes: Based on the received conversation data, using a pre-trained language model to take into account the context of each language symbol to extract each third feature vector of each language symbol; Based on the extracted third feature vectors, the first feature vector of the sentence is determined through comprehensive operation.

11. The conversation intention classification device according to claim 10, characterized in that: Determining the first feature vector of the sentence through comprehensive operations based on the extracted third feature vectors further includes: performing a pooling operation on each of the extracted third feature vectors to determine a fourth feature vector of the sentence; and Based on the fourth eigenvector of the sentence, a dense network layer is used for processing to obtain the first eigenvector of the sentence.

12. The dialogue intention classification device according to claim 11, characterized in that, The pre-trained language model includes a BERT language model, which includes 12-24 conversion layers, and only the last conversion layer and the dense network layer are trained during the joint training of the first learning network and the second learning network.

13. The dialogue intention classification device according to claim 9, characterized in that The representative feature vector of the known intent category is the cluster center vector of the known intent category, the difference between the first feature vector and the representative feature vector includes the distance between the first feature vector and the representative feature vector, and the second learning network includes a cosine classifier.

14. The dialogue intention classification device according to claim 13, characterized in that Based on the second feature vector, using the trained second learning network to determine the various probability-related parameters of each known intent category further includes: normalizing the second feature vector and normalizing the weight of the cosine classifier; based on the normalized second feature vector, using the cosine classifier after weight normalization to determine the various probability-related parameters of each known intent category.

15. The conversation intention classification device according to claim 13, characterized in that: The second loss function includes a cross entropy loss function, and the first loss function is defined using the following formula (1): Among them, Loss1 represents the first loss function or its first component, ||·||2 represents the Euclidean distance, i represents the sample serial number, N represents the total number of samples, K represents the number of known intention categories, k represents the known intention category serial number, and y i represents the known intention category label of the i-th sample and y i ∈{1,2,…,K}, m1 represents the set loss boundary, z i represents the first feature vector extracted for the i-th sample, and c k represents the cluster center vector of the k-th known intention category.

16. The dialogue intention classification device according to claim 15, characterized in that, The first loss function is defined by combining the components Loss1, Loss2, and Loss3 of the first loss function defined by formula (1), the following formula (2), and formula (3): where Loss2 represents the second component of the first loss function, and s1 is a scaling factor. is the first eigenvector z i and the weight vector i of the known intent category identified by the label y The included angle between them, j represents the known intent category serial number, and θ j is the first eigenvector z i and the weight vector W of the j-th category j The included angle between them, and m2 is the cosine distance margin constant. Among them, Loss3 represents the third component of the first loss function, s2 is the scaling factor, and m3 is the angular distance marginal constant.

17. A non-volatile computer storage medium having executable instructions stored thereon, wherein when the executable instructions are executed by a processor, a method for classifying conversational intent is implemented, comprising: Based on the received conversation data, using the trained first learning network, extracting a first feature vector of the sentence; Based on a minimum difference between the first feature vector and representative feature vectors of each known intent category, adjusting the modulus of the first feature vector so that the smaller the minimum difference, the larger the modulus, to obtain a second feature vector; Based on the second feature vector, determining probability-related parameters for each known intent category using a trained second learning network, where the second learning network is configured to perceive the modulus of the second feature vector, wherein the first learning network and the second learning network are jointly trained using a first loss function and a second loss function representing classification loss, wherein the first loss function is defined such that a difference between the first feature vector and a representative feature vector of the known intent category to which it belongs is smaller than representative values of differences between the first feature vector and representative feature vectors of other known intent categories, and the probability-related parameters are parameters related to the probability of dialog intent classification; Compare the probability-related parameters of each known intent category determined with the threshold. When all probability-related parameters are less than the threshold, it is determined that the conversation belongs to an unknown intent category. When all probability-related parameters are above the threshold, it is determined that the conversation belongs to the known intent category corresponding to the maximum probability-related parameter.

Citation Information

Patent Citations

  • Answer data generation method based on dialog system and relevant device

    CN108009287A

  • Semantic comprehension method in task type dialogue system

    CN111104498A