A Method and System for User Intention Recognition Based on Deep Active Learning

Through lightweight deep neural network and incremental training combined with multi-criteria selection strategies, the problems of scarcity of training data and redundant sample selection in deep learning models are solved, and efficient and accurate user intention recognition is achieved.

CN114741500BActive Publication Date: 2025-07-29INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110018869.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-07
Publication Date
2025-07-29
Estimated Expiration
2041-01-07

AI Technical Summary

Technical Problem

Deep learning models require a large amount of training data in user intention recognition, high labeling costs, and there are problems such as large amount of computation and redundant or isolated points in active learning.

Method used

Lightweight deep neural network structure and incremental training method are adopted, combined with multi-criteria selection strategies, and high-value samples are selected through information, representation and diversity to mark them, reducing the amount of calculation and improving sample quality.

Benefits of technology

It improves the efficiency and accuracy of user intention recognition, reduces the amount of calculation, avoids redundancy and isolated points of sample selection, and improves the overall performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114741500B_ABST
    Figure CN114741500B_ABST
Patent Text Reader

Abstract

The present invention discloses a user intention recognition method and system based on deep active learning. The steps of this method include: 1) The data preprocessing module preprocesses the text describing the user intention to obtain an unlabeled corpus U; 2) The classification module classifies and predicts the samples in the unlabeled corpus U, obtains the prediction probabilities of the samples and outputs them to the selection module; 3) The selection module selects the k samples with the highest value based on the prediction probabilities of the samples and the set multi-criterion selection strategy, labels them and adds them to the labeled corpus, and deletes the k samples from the unlabeled corpus U; then uses the updated labeled corpus to train and update the classification module; 4) Repeat steps 2 to 3) until the iteration termination condition is met to obtain the trained classification module; 5) Use the trained classification module to perform intention classification prediction on the text to be recognized to obtain the predicted user intention category.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for user intention recognition based on deep active learning, belonging to the field of software technology. Background Art

[0002] Accurately recognizing user intentions is a prerequisite for many applications such as semantic search, personalized recommendation, and intelligent question answering. Intention recognition aims to determine the true intention of a user through the language described by the user, which belongs to a classification problem. Currently, deep learning models have made great progress in text classification problems and achieved performance far beyond traditional methods. However, for deep learning models to achieve good performance, a large amount of training data is required. In practical applications, labeled training corpora are often scarce and difficult to obtain. Therefore, how to reduce the annotation cost of training data for deep learning models is crucial.

[0003] Active learning provides a method for assisting in annotating corpora. By a certain selection strategy, the value of unlabeled samples is calculated, and high-value samples that are more in need of annotation are actively selected and handed over to experts for annotation. However, there are still the following problems in using active learning to complete the corpus annotation for the intention recognition task: on the one hand, after each round of annotation in the active learning iteration process, the corpus needs to be updated and the model needs to be retrained, which is computationally expensive for complex deep learning models; on the other hand, the commonly used uncertainty-based selection strategy in active learning does not consider the value of samples in the corpus space, resulting in the selected samples being isolated points or redundant problems. Summary of the Invention

[0004] Aiming at the problems existing in the existing user intention recognition technology, the purpose of the present invention is to provide a method and system for user intention recognition based on deep active learning. The present invention adopts an active learning framework. On the one hand, it adopts a lightweight deep neural network structure and an incremental training method, which greatly reduces the computational amount of model training and improves the system execution efficiency. On the other hand, it adopts a multi-criterion selection strategy, and selects high-value samples based on multiple criteria such as sample informativeness, representativeness, and diversity and hands them over to experts for annotation, increasing the quantity and quality of annotated samples and improving the overall performance of intention recognition. The present invention not only considers the information of individual samples, but also considers the overall distribution of the sample space, avoiding the problems of isolated points and redundancy of the selected samples to be annotated.

[0005] Technical solution of the present invention: A user intention recognition method and system based on deep active learning combines deep learning with active learning. Deep learning is used to solve the problem of user intention classification, and active learning is used to solve the problem of scarce labeled corpus faced in the training process of the deep learning model. In the system, aiming at the problem of large computational cost in iterative training of the deep learning model on an ever-updating corpus, a lightweight deep neural network structure is designed, and an incremental training method is designed. Before selecting a new round of samples, the network weights are updated with a small number of iterations; aiming at the problem of the strategy for selecting samples to be labeled, a multi-criterion selection strategy based on sample informativeness, representativeness, and diversity is adopted to avoid the problems of outliers and redundancy in the selected samples to be labeled. The system includes a data preprocessing module, a classification module, a selection module, an iteration controller, and a labeled corpus. Among them:

[0006] The data preprocessing module performs preprocessing operations such as character cleaning and word segmentation on the text describing the user intention. Its input is the unlabeled original corpus, that is, the text describing the user intention; the output is the unlabeled corpus set U after the above preprocessing operations.

[0007] The classification module performs intention classification prediction on the text describing the user intention. Its core is a two-layer CNN classification model (refer to Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. ImageNet classification with deep convolutional neural networks. In Proceedings of the 25th International Conference on Neural Information Processing Systems-Volume 1 (NIPS'12). Curran Associates Inc., Red Hook, NY, USA, 1097–1105.). In the model training stage, the samples in U are classified and predicted to obtain the prediction probability P of the samples, which is output to the selection module, and the model parameters are fine-tuned based on the updated corpus after the labeled corpus is updated to update the classification model; in the model classification stage, intention classification prediction is performed on the text describing the user intention. The input is the text describing the user intention, and the output is the predicted intention category. The classification module includes the following three sub-modules:

[0008] (1) Embedding representation sub-module: It performs word-level feature representation on the preprocessed user intention text, and each word corresponds to a vector with a fixed dimension. This vector consists of two parts: a word vector and a position vector. The word vector represents the semantic features of the word and is pre-trained using the Word2Vec method on the Chinese Wikipedia corpus. The position vector represents the position information and relative distance features of the word in the text, which is randomly initialized and then obtained through model training.

[0009] (2) Feature extraction sub-module: It performs sentence-level feature representation on the user intention text. Sentence features include context information, sentence structure, sentence semantics, etc. This sub-module consists of a two-layer convolutional CNN network. After each convolution, batch normalization and a rectified linear unit are added to avoid gradient vanishing and accelerate the convergence of model training. The output of each convolutional kernel is represented by the concatenation of max pooling and average pooling. Finally, the results of all convolutional kernels are concatenated to obtain the representation vector of the user intention text.

[0010] (3) Output sub-module: It outputs the probability that the user intention text belongs to each intention category. This sub-module consists of a two-layer fully connected network, Dropout, and Softmax. Among them, Dropout is used to alleviate the overfitting problem of the model, and Softmax normalizes to obtain the probability that the user intention text belongs to each intention category.

[0011] The data processing process of the classification module is as follows:

[0012] (1) The preprocessed user intention text is input into the embedding representation sub-module to obtain a set of word-level vectors.

[0013] (2) The above word-level vectors are input into the feature extraction sub-module. First, they go through two-layer convolutional operations, where batch normalization and a rectified linear unit are added after each convolution, and then through a pooling operation, using max pooling and concatenation representation. Finally, the results of all convolutional kernels are concatenated to obtain the representation vector of the user intention text.

[0014] (3) The above user intention text vector is input into the output sub-module, going through a two-layer fully connected network, Dropout, and Softmax, to obtain the probability that the user intention text belongs to each intention category.

[0015] Selection module: Based on the prediction probability P of the samples output by the classification model and the multi-criteria selection strategy, it calculates the value of the samples, outputs the k samples with the highest value, and hands them over to the experts for manual annotation. At the same time, these k samples are deleted from the unlabeled corpus U. The implementation steps of this module are as follows:

[0016] (1) Calculate the sample information amount based on information entropy. The formula is as follows:

[0017]

[0018] Among them, U = {x1…x n} is the unlabeled corpus, xi is the i-th unlabeled corpus sample in the unlabeled corpus U, Y represents all categories, i and Pr(xi, r) represents the probability that the sample xi is predicted as the r-th (r ∈ [1, |Y|]) category by the model. Pr(xi, r) represents the probability that the sample xi i is predicted as the r-th (r ∈ [1, |Y|]) category by the model.

[0019] (2) Calculate the sample density based on the sample similarity and use it as a measure of the representativeness of the sample in the unlabeled corpus. The similarity calculation formula between the sample xi i and the sample xj j is as follows:

[0020]

[0021] Among them, U = {x1…x n} is the unlabeled dataset, Y represents all categories, Pr(xi, r) and Pr(xj, r) respectively represent the probability that the samples xi i , xj j are predicted as the r-th (r ∈ [1, |Y|]) category by the model.

[0022] The density of the sample xi i in the unlabeled corpus U is defined as the average of the similarities between xi i and all other samples in the corpus U, and its calculation formula is as follows:

[0023]

[0024] where n is the total number of samples in the unlabeled corpus U.

[0025] (3) Calculate the sample value in combination with the sample diversity and select the k samples with the highest value. The steps are as follows:

[0026] (3.1) Use the k-Means algorithm to cluster the samples in the unlabeled corpus U into k classes.

[0027] (3.2) Traverse the k classes in turn. For each class, calculate the value of each sample in it and select the sample with the highest value in this class. The formula is as follows:

[0028]

[0029] where λ is a tuning parameter and 0 < λ < 1.

[0030] (3.3) Output the samples with the highest value in the k classes in turn, a total of k samples.

[0031] Iterative controller, which controls the number of iterations of model training. When the iterative termination condition is met, the model training ends; otherwise, the operations of the classification module and the selection module on the unlabeled corpus U are repeated.

[0032] Labeled corpus, which stores the labeled corpus. The k samples selected by the selection module are stored in the labeled corpus after being labeled by experts, and the labeled corpus completes one round of update.

[0033] The advantages of the present invention compared with the prior art are as follows:

[0034] (1) The system adopts a lightweight double-layer CNN network structure, which can well support parallelization. At the same time, an incremental training method is adopted to mix the newly added labeled samples with the original samples, and before the new round of sample selection, the network weights are updated with a small number of iterations, thus greatly reducing the computational amount of model training and improving the system execution efficiency.

[0035] (2) When selecting samples for annotation, a selection strategy combining multiple criteria is adopted, and the sample value is calculated by comprehensively considering the informativeness, representativeness and diversity of the samples. This strategy not only considers the information of individual samples, but also considers the overall distribution of the sample space, avoiding the problems of isolated points and redundancy in the selected samples to be annotated, which helps to improve the quality of the annotated samples and thus improve the overall performance of the system. Description of the Drawings

[0036] Figure 1 It is the system architecture diagram of the present invention.

[0037] Figure 2 It is the model training process diagram of the present invention system. Detailed Embodiments

[0038] The present invention will be described in detail below with reference to specific examples.

[0039] As Figure 1 shown, the system of the present invention includes five major modules: a data preprocessing module, a classification module, a selection module, an iterative controller, and a labeled corpus. Among them, the labeled corpus is used to store the labeled corpus. The iterative controller is used to control the number of iterations of model training. When the iterative termination condition is met, the model training ends; otherwise, operations such as iterative loop classification prediction, selection of unlabeled corpus, update of the corpus, and update of the classification model are performed. The data preprocessing module is used to perform preprocessing operations such as character cleaning and word segmentation on the text describing the user's intention. For example, the original corpus with the input of "How to sign up for the exam?"; the output is the preprocessed unlabeled corpus x i=(How to register for an exam). The classification module is used to perform intent classification prediction on the text describing the user's intent. In the model training stage, the predicted classification probability is output to the selection model, and the model parameters are fine-tuned after the labeled corpus is updated; in the model classification stage, the predicted classification result is directly output. The selection module is used to calculate the sample value and output the k samples with the highest value as the samples to be labeled.

[0040] The classification module contains three sub-modules. Among them, the embedding representation sub-module is used to perform word-level feature representation on the preprocessed unlabeled corpus. Each word corresponds to a vector of a fixed dimension, which includes two parts: a word vector and a position vector, and their dimensions can be set to 200 and 100 respectively. The word vector is pre-trained using the Word2Vec method in the Chinese Wikipedia corpus, and the position vector is randomly initialized and then obtained through model training. The feature extraction sub-module consists of a two-layer CNN network and is used to perform sentence-level feature representation on the user intent text. After each convolution, batch normalization and a rectified linear unit are added, and the output of each convolution kernel is represented by the concatenation of max pooling and average pooling. The results of all convolution kernels are concatenated to obtain the representation vector of the user intent text. Among them, the convolution kernel scale is set to (1, 2, 3, 4), the number of channels is 128, the batch size is 128, and the model training learning rate is 0.001. The output sub-module consists of a two-layer fully connected network, Dropout, and Softmax, and is used to output the probability that the user intent text belongs to each intent category. Among them, the dropout probability is 0.5.

[0041] The model training process in the system of the present invention is as Figure 2 shown. First, the data preprocessing module performs preprocessing operations such as character cleaning and word segmentation on the text describing the user's intent to obtain an unlabeled corpus set U. Then, it enters the model iterative training process. In this process, first, the classification model predicts the samples in U to obtain the sample prediction probability; then, the selection module calculates the sample value according to the multi-criterion selection strategy and selects the k samples with the highest value to be labeled by an expert. At the same time, these k samples are deleted from U; then, the k labeled samples are stored in the labeled corpus, and the classification model is fine-tuned on the updated labeled corpus; finally, the control iterator determines whether the preset end condition is reached. If it is reached, the model training process ends; otherwise, the above model training process is iteratively looped.

[0042] Among them, the process of the selection module calculating the sample value according to the multi-criterion selection strategy and selecting the Top(k) high-value samples is as follows: First, calculate the information amount of the sample based on the information entropy of the prediction probabilities of all classes for the sample; then, calculate the sample density based on the sample similarity and use it as a measure of the representativeness of the sample in the unlabeled corpus; finally, combined with the sample diversity, perform k-Means clustering on the unlabeled corpus, and within each class, calculate the sample value based on the information amount and representativeness measure of the sample, and select the sample with the highest value within the class. Eventually, k samples with the highest value are selected from the k classes.

[0043] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the present invention without departing from the principles and scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.

Claims

1. A method for user intention recognition based on deep active learning, the steps of which include: 1) The data preprocessing module preprocesses the text describing the user intention to obtain an unlabeled corpus U; 2) The classification module classifies and predicts the samples in the unlabeled corpus U, obtains the prediction probabilities of the samples and outputs them to the selection module; the classification module includes an embedding representation sub-module, a feature extraction sub-module and an output sub-module; wherein, the embedding representation sub-module is used to perform word-level feature representation on the samples in the unlabeled corpus U and input them into the feature extraction sub-module, where each word corresponds to a vector of a fixed dimension, and this vector includes two parts: a word vector representing the semantic features of the word and a position vector, and this position vector represents the position information and relative distance features of the word in the text; the feature extraction sub-module is a double-layer CNN classification model, which is used to process the input information to obtain a user intention text representation vector and input it into the output sub-module; the output sub-module is used to process the input information to obtain the probability that the user intention text belongs to each intention category; 3) The selection module selects the top k samples with the highest value based on the predicted probabilities of the samples and the set multi-criteria selection strategy, annotates them, and adds them to the annotated corpus, and deletes the k samples from the unannotated corpus U; then uses the updated annotated corpus to train and update the classification module; wherein, the method for selecting the top k samples with the highest value based on the predicted probabilities of the samples and the set multi-criteria selection strategy is: First, calculate the sample information amount Info(x i of the sample x i ), calculate the representative measure Repr(x i ) of the sample x i based on sample similarity; then cluster the samples in the unannotated corpus U into k classes; then traverse the k classes in turn, and for each class, calculate the value of each sample therein and select the sample with the highest value in the class; the sample information amount i of the sample x x i is the i-th unannotated sample in the unannotated corpus U, Y represents all classes, represents the probability that the sample x i is predicted to be the r-th class; the representative measure i of the sample x n is the total number of samples in the unannotated corpus U, Y represents all classes, respectively represent the probabilities that the samples x i and x j in the unannotated corpus U are predicted to be the r-th class; 4) Repeat steps 2 to 3) until the iteration termination condition is satisfied to obtain a trained classification module; 5) Use the trained classification module to perform intention classification prediction on the text to be recognized to obtain the predicted user intention category.

2. The method according to claim 1, characterized in that, Calculate the value of sample x using the formula λInfo(x i )+(1-λ)Repr(x i ); where λ is the weight coefficient. i ​ 3. A user intention recognition system based on deep active learning, characterized in that, It includes a data preprocessing module, a classification module, and a selection module, where the data preprocessing module is used to preprocess the text describing the user intention to obtain an unlabeled corpus U; the classification module is used to classify and predict the samples in the unlabeled corpus U to obtain the prediction probabilities of the samples; the classification module includes an embedding representation sub-module, a feature extraction sub-module and an output sub-module; wherein, the embedding representation sub-module is used to perform word-level feature representation on the samples in the unlabeled corpus U and input them into the feature extraction sub-module, where each word corresponds to a vector of a fixed dimension, and this vector includes two parts: a word vector representing the semantic features of the word and a position vector, and this position vector represents the position information and relative distance features of the word in the text; the feature extraction sub-module is a double-layer CNN classification model, which is used to process the input information to obtain a user intention text representation vector and input it into the output sub-module; the output sub-module is used to process the input information to obtain the probability that the user intention text belongs to each intention category; The selection module is used to select the top k samples with the highest value based on the prediction probability of the samples and the set multi-criteria selection strategy, label them, and add them to the labeled corpus, and delete the k samples from the unlabeled corpus U; among them, the selection module first calculates the sample information amount Info(x i ) of the sample x i , calculates the representative measure Repr(x i ) of the sample x i based on the sample similarity; then clusters the samples in the unlabeled corpus U into k classes; then traverses the k classes in turn, and for each class, calculates the value of each sample in it and selects the sample with the highest value in the class; the sample information amount i of the sample x x i is the i-th unlabeled sample in the unlabeled corpus U, Y represents all classes, represents the probability that the sample x i is predicted to be the r-th class; the representative measure i of the sample x n is the total number of samples in the unlabeled corpus U, Y represents all classes, respectively represent the probabilities that the samples x i and x j in the unlabeled corpus U are predicted to be the r-th class; wherein in the training stage of the classification module, the updated labeled corpus is used to train and update the classification module until the iteration termination condition is satisfied; in the prediction stage, the trained classification module is used to perform intention classification prediction on the text to be recognized to obtain the predicted user intention category.