Method for generating user intention recognition model, user intention recognition method and device

Through layer-by-layer knowledge distillation, search for micro-network structures and fine-tuning of models, the user intention recognition model is automatically obtained, which solves the problem of automatically obtaining small models in the existing technology, and realizes fully automatic intelligent acceleration of the identification of user intentions and model performance improvement.

CN113988267BActive Publication Date: 2025-06-10CTRIP TRAVEL INFORMATION TECH (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111295501.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-03
Publication Date
2025-06-10
Estimated Expiration
2041-11-03

AI Technical Summary

Technical Problem

The prior art is difficult to obtain small models suitable for user intention recognition automatically, resulting in the need of a large amount of human resources and hardware resources for model parameter adjustment and training.

Method used

Through layer-by-layer knowledge distillation, micro-neural network structure search and model fine-tuning, the user's intention to identify the model is automatically obtained, reducing the amount of model parameters and hardware resource consumption.

Benefits of technology

It realizes fully automatic intelligent acceleration to identify user intentions, reduces model size and reasoning time, while improving model performance, reducing manual intervention and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113988267B_ABST
    Figure CN113988267B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and provides a method for generating a user intention recognition model, a user intention recognition method and a device. The generation method includes: training a teacher model including multiple layers of encoding networks based on layer-by-layer knowledge distillation to obtain the output Logits of each layer of encoding network of the target teacher model; based on a student model including multiple layers of convolutional networks, performing differentiable neural network architecture search according to a target loss function including the cross-entropy loss between the output result of the student model and the true label of the training data, and the cross-entropy loss between the output results of each layer of convolutional networks and the output Logits of each layer of encoding networks, to obtain a target student model; and fine-tuning the target student model according to the cross-entropy loss between the output result of the target student model and the true label to obtain a user intention recognition model. The present invention automatically obtains a user intention recognition model through knowledge distillation, differentiable search and fine-tuning, reducing the number of model parameters and the consumption of hardware resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and more specifically, to a method for generating a user intention recognition model, a user intention recognition method, and a device. Background Art

[0002] Large online travel service platforms can provide one-stop professional customer service reservation services such as hotels, air tickets, train tickets, routes, tickets for entertainment, visas, and corporate business travel to a vast number of users. Facing a huge number of users, in order to reduce the cost of artificial customer service, it is necessary to effectively identify the content input by users, determine whether it is a business consultation or just casual chat. If it is just casual chat, it will be directly transferred to an intelligent customer service, that is, a customer service robot. How to quickly and accurately identify user intentions is the focus of attention of online travel service platforms.

[0003] Currently, the BERT (Bidirectional Encoder Representations from Transformer) model has been proven to be effective in various NLP (Natural Language Processing) tasks. However, the BERT model generally has a large number of parameters and a huge model size, resulting in difficulty in training and applying the BERT model. It is necessary to study how to reduce the BERT model and accelerate the model inference speed.

[0004] There are many existing methods dedicated to compressing the BERT model into a small model and accelerating its inference speed. The current methods mainly include: distilling the logits, hidden parameters, embeddings, etc. of the BERT model to transfer the capabilities of the BERT model to a small model; quantizing the BERT model and using low precision or mixed precision to reduce the model size; pruning redundant parameters from the BERT model.

[0005] Although the above methods can reduce the model parameters and accelerate the model inference speed, they all require engineers to continuously adjust parameters and conduct cumbersome experiments such as training, which will consume a large amount of human and hardware resources.

[0006] Therefore, how to automatically obtain a small model suitable for user intention recognition and achieve fully automatic acceleration of user intention recognition is an urgent problem to be solved in this field.

[0007] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present invention, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0008] In view of this, the present invention provides a method for generating a user intention recognition model, a user intention recognition method and a device, which can automatically obtain a user intention recognition model through layer-by-layer knowledge distillation, differentiable neural network architecture search and model fine-tuning, reduce the number of model parameters and the consumption of hardware resources, and achieve fully automatic intelligent acceleration for recognizing user intentions.

[0009] One aspect of the present invention provides a method for generating a user intention recognition model, including: training a teacher model including multiple layers of encoding networks based on layer-by-layer knowledge distillation to obtain the output Logits of each layer of encoding networks of the target teacher model; based on a student model including multiple layers of convolutional networks, performing differentiable neural network architecture search according to an objective loss function including the cross-entropy loss between the output result of the student model and the true label of the training data, and the cross-entropy loss between the output results of each layer of the convolutional networks and the output Logits of each layer of the encoding networks, to obtain a target student model; and fine-tuning the target student model according to the cross-entropy loss between the output result of the target student model and the true label, to obtain a user intention recognition model.

[0010] In some embodiments, the training of the teacher model including multiple layers of encoding networks based on layer-by-layer knowledge distillation includes: inserting a Probe classifier corresponding to each layer of the encoding networks in the teacher model; and training the teacher model based on knowledge distillation by using the training data, so as to obtain the output Logits of each layer of the encoding networks through the Probe classifier.

[0011] In some embodiments, the obtaining of the output Logits of each layer of encoding networks of the target teacher model includes: training the teacher model for a preset number of times; and obtaining the optimal output of each layer of the encoding networks in the preset number of times of training as the output Logits of each layer of the encoding networks of the target teacher model.

[0012] In some embodiments, the teacher model adopts a Bidirectional Encoder Representations from Transformers (BERT) model; the candidate operators of each layer of the convolutional networks include: multiple convolutions with different convolution kernel sizes, multiple dilated convolutions with different convolution kernel sizes, average pooling, max pooling, Identity function, and Zero function.

[0013] In some embodiments, when performing differentiable neural network architecture search, each convolutional network layer is used as a search unit. Each search unit includes two input nodes, one output node, and multiple intermediate nodes. In each search unit, the two input nodes are the output nodes of the previous two search units. Each intermediate node is connected to the output node, and each intermediate node has two incoming edges, with each incoming edge selected from the candidate operators.

[0014] In some embodiments, the target loss function further includes an efficiency-aware loss for differentiable neural network architecture search. The formula of the target loss function is:

[0015]

[0016] where is the cross-entropy loss between the output result of the student model and the true label, is the cross-entropy loss between the output results of each convolutional network layer and the output Logits of each encoding network layer, is the efficiency-aware loss, and γ and β are hyperparameters.

[0017] In some embodiments, the cross-entropy loss between the output result of the i-th convolutional network layer and the output Logits of the j-th encoding network layer

[0018]

[0019] where is the Probe classifier on the j-th encoding network layer, is the hidden representation of the j-th encoding network layer, is the output Logits of the j-th encoding network layer, is the Probe classifier on the i-th convolutional network layer, is the hidden representation of the i-th convolutional network layer, is the output result of the i-th convolutional network layer, and T is the temperature coefficient;

[0020]

[0021] where M is the number of samples in the training data, K is the number of convolutional network layers in the student model, ω i,m is the normalized weight of the cross-entropy loss ;

[0022]

[0023] where y mis the label of the m-th sample, where the positive class is 1 and the negative class is 0.

[0024] In some embodiments, the efficiency-aware loss has the following formula:

[0025]

[0026] where K is the number of convolutional network layers of the student model, and K max is the predefined maximum number of layers, o i,j is the candidate operator to be searched, α c is the network structure of the target student model, SIZE(.) is the parameter size, and FLOPs(.) is the number of floating-point operations for searching each candidate operator.

[0027] In some embodiments, the cross-entropy loss between the output result of the student model and the true label has the following formula:

[0028]

[0029] where y m is the label of the m-th sample, where the positive class is 1 and the negative class is 0, and p(y m ) is the probability that the student model predicts the m-th sample as the positive class; when fine-tuning the target student model, p(y m ) is the probability that the target student model predicts the m-th sample as the positive class.

[0030] In some embodiments, the training data includes chitchat samples and non-chitchat samples; the user intent recognition model is used to identify whether the user input content is chitchat.

[0031] Another aspect of the present invention provides a user intent recognition method, including: obtaining user input content; inputting the user input content into the user intent recognition model generated by the generation method described in any of the above embodiments to obtain a user intent recognition result; and based on the user intent recognition result, determining whether the user input content is chitchat. If it is, transfer to an intelligent customer service; if not, transfer to a human customer service.

[0032] Another aspect of the present invention provides a device for generating a user intention recognition model, including: a knowledge distillation module, configured to train a teacher model including a multi-layer encoding network based on layer-by-layer knowledge distillation to obtain the output Logits of each layer of the encoding network of the target teacher model; a differentiable search module, configured to perform differentiable neural network architecture search based on a student model including a multi-layer convolutional network according to an objective loss function including the cross-entropy loss between the output result of the student model and the true label of the training data, and the cross-entropy loss between the output results of each layer of the convolutional network and the output Logits of each layer of the encoding network, to obtain a target student model; and a fine-tuning processing module, configured to fine-tune the target student model according to the cross-entropy loss between the output result of the target student model and the true label to obtain a user intention recognition model.

[0033] Another aspect of the present invention provides a user intention recognition device, including: a content acquisition module, configured to acquire user input content; an intention recognition module, configured to input the user input content into a user intention recognition model generated by the generation method according to any of the above embodiments to obtain a user intention recognition result; and a transfer processing module, configured to determine whether the user input content is a chat according to the user intention recognition result, and if so, transfer to an intelligent customer service, and if not, transfer to a human customer service.

[0034] Another aspect of the present invention provides an electronic device, including: a processor; and a memory storing executable instructions therein; wherein, when the executable instructions are executed by the processor, the generation method of the user intention recognition model according to any of the above embodiments is implemented, and / or, the user intention recognition method according to any of the above embodiments is implemented.

[0035] Another aspect of the present invention provides a computer-readable storage medium for storing a program, which, when executed by a processor, implements the generation method of the user intention recognition model according to any of the above embodiments, and / or, implements the user intention recognition method according to any of the above embodiments.

[0036] The beneficial effects of the present invention compared with the prior art at least include:

[0037] Through layer-by-layer knowledge distillation and differentiable neural network architecture search, a target student model with reduced model parameter quantity and model size is automatically obtained, and at the same time, the consumption of hardware resources can be reduced; after the model architecture search is completed, fine-tuning is performed to significantly improve the model performance without re-training, and a user intention recognition model is obtained, realizing fully automatic intelligent acceleration for recognizing user intentions.

[0038] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention. Obviously, the accompanying drawings described below are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0040] Figure 1 Schematic diagram showing the steps of a method for generating a user intention recognition model in an embodiment of the present invention;

[0041] Figure 2 Schematic diagram of a neural network architecture for performing differentiable neural network architecture search in an embodiment of the present invention;

[0042] Figure 3 Schematic diagram showing the curve relationship between the number of update iterations and the target loss function of differentiable neural network architecture search in an embodiment of the present invention;

[0043] Figure 4 Schematic diagram showing the curve relationship between the number of update iterations and the accuracy of the student model of differentiable neural network architecture search in an embodiment of the present invention;

[0044] Figure 5 Schematic diagram showing the steps of a user intention recognition method in an embodiment of the present invention;

[0045] Figure 6 Schematic diagram of the modules of a device for generating a user intention recognition model in an embodiment of the present invention;

[0046] Figure 7 Schematic diagram of the modules of a user intention recognition device in an embodiment of the present invention;

[0047] Figure 8 Schematic diagram of the structure of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art.

[0049] The accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0050] In addition, the processes shown in the accompanying drawings are only exemplary illustrations and do not necessarily include all steps. For example, some steps can be decomposed, some steps can be combined or partially combined, and the actual execution order may change according to the actual situation. The terms "first", "second" and similar terms used in the specific description do not denote any order, quantity or importance, but are only used to distinguish different components. It should be noted that, without conflict, the embodiments of the present invention and the features in different embodiments can be combined with each other.

[0051] The online travel service platform has a large number of users and provides users with real-time online customer service consultation services. In order to reduce the cost of artificial customer service, it is necessary to quickly identify the intention of the user input content and classify it, that is, whether it is chatting or business consultation, and then transfer it to the intelligent customer service robot or artificial customer service accordingly. User intention recognition, as a text classification task of NLP tasks, generally uses the BERT model. However, the BERT model is huge, has a large number of parameters, and has a slow inference speed, which will cause a certain delay in identifying the user's intention. In order to reduce the number of parameters of the BERT model, accelerate the model inference speed while reducing the consumption of engineers' and hardware resources, the present invention proposes a model optimization method for user intention recognition tasks, which can automatically accelerate user intention recognition and avoid the cumbersome work of manually pruning and quantifying the BERT model.

[0052] The present invention realizes automatic acceleration by adaptively searching for a small model for user intention, and uses differentiable neural network architecture search to automatically compress the large BERT model (teacher model) into a small model (student model) suitable for user intention recognition, which not only reduces the model parameters but also speeds up the model inference speed. Using knowledge distillation to provide clues for the search of the small model and constraining FLOPs and Size, so that the search model can have a balance between efficiency and effectiveness. After obtaining the small model, the present invention also fine-tunes the search model to further improve the model performance without re-training.

[0053] Figure 1 Shows the main steps of the method for generating a user intention recognition model in an embodiment, referring to Figure 1 As shown, the method for generating a user intention recognition model includes:

[0054] Step S110: Based on layer-by-layer knowledge distillation, train a teacher model including a multi-layer encoding network to obtain the output Logits of each layer of the encoding network of the target teacher model.

[0055] The teacher model can adopt the BERT model. The BERT model is a pre-trained model. In this embodiment, a 12-layer BERT model can be used. Additionally, in this embodiment, the user intention recognition model is used to identify whether the input content of the user on the online travel service platform belongs to casual chat. Therefore, the training data for training the teacher model includes casual chat samples and non-casual chat samples. For example, in the training data, there are 3033 casual chat samples and 4113 non-casual chat samples. The 3033 casual chat samples are labeled as 1, and the 4113 non-casual chat samples are labeled as 0.

[0056] In other embodiments, BERT models with other numbers of layers can be adopted, such as 6-layer, 24-layer, and so on. Additionally, in other embodiments, if the user intention recognition model is used to identify other user intentions, such as identifying the business type that the user needs to consult, then user input samples including various business types can be used as training data.

[0057] Training a teacher model including a multi-layer encoding network based on layer-by-layer knowledge distillation specifically includes: inserting a Probe classifier corresponding to each layer of the encoding network in the teacher model; using the training data to train the teacher model based on knowledge distillation to obtain the output Logits of each layer of the encoding network through the Probe classifier.

[0058] For a large neural network trained on a main task, a Probe (probe) is a shallow neural network inserted in its intermediate hidden layer, usually a classifier layer. By training the teacher model, each layer of the Probe classifier of the teacher model learns to achieve layer-by-layer knowledge distillation using the probes and obtain the output Logits of each layer of the encoding network of the optimal teacher model.

[0059] Obtaining the output Logits of each layer of the encoding network of the target teacher model specifically includes: training the teacher model for a preset number of times; obtaining the optimal output of each layer of the encoding network in the preset number of trainings as the output Logits of each layer of the encoding network of the target teacher model.

[0060] Step S120: Based on a student model including a multi-layer convolutional network, perform differentiable neural network structure search according to a target loss function including the cross-entropy loss between the output result of the student model and the true label of the training data, and the cross-entropy loss between the output results of each layer of the convolutional network and the output Logits of each layer of the encoding network to obtain the target student model.

[0061] The student model is predefined and includes multiple convolutional networks. The candidate operators for each convolutional network layer include: multiple convolutions with different kernel sizes, such as cnn3 (3*3 kernel), cnn5 (5*5 kernel), and cnn7 (7*7 kernel); multiple dilated convolutions with different kernel sizes, such as dilated_cnn3, dilated_cnn5, and dilated_cnn7; average pooling avg_pool, max pooling max_pool, Identity function, and Zero function.

[0062] Figure 2 Figure 1 shows a neural network architecture for differentiable neural network architecture search in one embodiment. Differentiable neural network architecture search, hereinafter referred to as DARTS (Differentiable Architecture Search), is performed. During the DARTS search process, first define Figure 2 the neural network architecture shown in Figure 2. In the student model 200, it includes an input layer Input, K convolutional network layers, and an output layer Classifier. When performing DARTS search, each convolutional network layer is used as a search unit Cell. Each search unit Cell includes two input nodes (node 0 and node 1), an output node (node 5, c_k), and multiple intermediate nodes (the number of intermediate nodes is optional. In this embodiment, 3 intermediate nodes are defined, namely node 2, node 3, and node 4). In each search unit Cell, the two input nodes are the output nodes of the previous two search units (node 0 is the output node c_{k - 2} of the previous previous search unit, and node 1 is the output node c_{k - 1} of the previous search unit). Each intermediate node is connected to the output node, and each intermediate node has two incoming edges, and each incoming edge is selected from the candidate operators. That is, in this embodiment, each edge entering the intermediate node has 10 candidate operators, namely: cnn3, cnn5, cnn7, dilated_cnn3, dilated_cnn5, dilated_cnn7, avg_pool, max_pool, identity, and zero. The purpose of DARTS search is to learn the architecture parameters α, indicating which operator should be selected for each edge and which edges should be retained.

[0063] In one embodiment, the objective loss function further includes an efficiency-aware loss for DARTS search. The formula of the objective loss function is: where is the cross-entropy loss between the output result of the student model and the true label, is the cross-entropy loss between the output results of each convolutional network layer and the output Logits of each layer of the encoding network, is the efficiency-aware loss, and γ and β are hyperparameters used to balance the three losses.

[0064] The target loss function can be specifically expressed as: where ω α is the trainable network weight of architecture α, is the loss of the target task, that is, the cross-entropy loss with respect to the true labels in the training data D t in the cross-entropy loss, is the knowledge distillation loss for the target task, which provides guidance for finding a model structure suitable for the target task, is the efficiency-aware loss, which provides restrictions on the model size and the number of floating-point operations FLOPs to help search for a lightweight and efficient structure. Through DARTS search, the three loss functions (cross-entropy supervised by true labels, cross-entropy supervised by layer-by-layer Probes, and structural regularization term) are optimized to obtain the best structure, that is, the target student model.

[0065] During the BERT model compression process, the Probe classifier is used to decompose useful task knowledge layer by layer from the teacher model, and then the knowledge is extracted into the compressed model. Specifically, first freeze the parameters of the teacher model, and train a Softmax probe classifier for each hidden layer (i.e., the encoding network) according to the target task labels. There are a total of J probe classifiers, and the classification logits of the j-th probe classifier are regarded as the knowledge learned from the j-th layer encoding network. Given the input instance m, let be represented as the hidden representation of the j-th layer encoding network of the target teacher model, and let be represented as the hidden representation of the i-th layer convolutional network of the student model. The useful task knowledge (output logits) extracted, that is, the cross-entropy loss between the output result of the i-th layer convolutional network and the output Logits of the j-th layer encoding network has the formula:

[0066]

[0067] where is the Probe classifier on the j-th layer encoding network, is the output Logits of the j-th layer encoding network, is the trainable Probe classifier on the i-th layer convolutional network, is the output result of the i-th layer convolutional network, and T is the temperature coefficient.

[0068] For different downstream tasks, since the role of each layer encoding network in the BERT model is different, the decomposed knowledge of all hidden layers is combined collectively as:

[0069]

[0070]

[0071] Among them, M is the number of samples of the training data, y m is the label of the m-th sample, with the positive class being 1 and the negative class being 0. K is the number of convolutional network layers of the student model, that is, the number of DARTS search stacking layers. ω i,m is the normalized weight of the cross-entropy loss That is, the normalized weight set for each sample according to the negative cross-entropy loss (loss) of the teacher model probe. The motivation for normalizing with loss is that for the layers with more accurate prediction results of the teacher model (smaller loss), the better they are trained or the more suitable they are for the target task, the higher weight will be assigned.

[0072] In order to obtain an effective target student model according to the BERT model, model efficiency is also incorporated into the loss function. It mainly includes parameter size and inference time. Specifically, for the searched structure α and the number of search stacking layers K, the efficiency-aware loss is defined as:

[0073]

[0074] Among them, K max is the predefined maximum number of layers, o i,j is the candidate operator to be searched, α c is the network structure of the target student model, SIZE(.) is the parameter size, and FLOPs(.) is the number of floating-point operations for searching each candidate operator.

[0075] The standard cross-entropy formula is adopted:

[0076]

[0077] Among them, p(ym) is the probability that the student model predicts the m-th sample as the positive class.

[0078] Figure 3 Shows the curve relationship between the update iteration times of DARTS search and the target loss function in an embodiment, Figure 4 Shows the curve relationship between the update iteration times of DARTS search and the accuracy of the student model. Referring to Figure 3 the curve 300 shown in Figure 4 and the curve 400 shown in

[0079] Step S130: Fine-tune the target student model according to the cross-entropy loss between the output result of the target student model and the true label to obtain a user intention recognition model.

[0080] When fine-tuning the target student model (Finetune), use the above formula, and replace p(y m ) with the probability that the target student model predicts the m-th sample as the positive class.

[0081] Directly using data and labels for fine-tuning through cross-entropy can further improve the model effect without the need for retraining.

[0082] For the user intention recognition task, that is, to determine whether the user input content is chitchat, after obtaining the user intention recognition model, use the validation set to conduct a validation experiment on the user intention recognition model. In the validation set, there are 808 chitchat sample data and 996 non-chitchat sample data. In the validation experiment, in order to verify the effectiveness of the user intention recognition model for the user intention recognition task, comprehensively compare the user intention recognition model with other mainstream BERT algorithms, including Accuracy, Precision, Recall, F1 value, as well as model size and speed. The results of the validation experiment are shown in the following table:

[0083] BERT BERT(*) BERT(**) ADABERT Ration Accuracy 0.93400 0.92794 0.95843 0.92905 96.9% Precision 0.93460 0.89882 0.94859 0.89813 94.7% Recall 0.93200 0.94554 0.95916 0.94926 99.0% F1 0.93310 0.92159 0.95385 0.92298 96.8% Size 110M 110M 110M 11M 10X Speed(CPU) 115ms 139ms 4.6ms 30X Speed(GPU) 12.9ms 4.5ms 3X

[0084] As can be seen from the above table, the user intention recognition model obtained by the present invention based on the optimized Adabert algorithm has similar performance to other mainstream algorithms, but the model size is reduced by about 10 times, and the model inference speed is significantly improved. The inference speed on the CPU is about 30 times that of BERT(*), and the inference speed on the GPU is about 3 times that of other mainstream algorithms, which proves the effectiveness of the present invention for accelerating the user intention recognition task.

[0085] In summary, the method for generating a user intention recognition model of the present invention can automatically search for a compressed model for the task of recognizing user intentions and perform fine-tuning to delete the task-specific redundant parts in the original large model, so as to achieve better compression of the model and faster inference speed while ensuring the model performance. The present invention regards the structure of the compressed small model as a set of learnable parameters, so that the algorithm can automatically search for the model structure to be used after compression for different tasks. The present invention places the compressed small model from the perspective of knowledge distillation and integrates the search target into the loss function of the student model. And this task is decomposed into two aspects. One is that the compressed model predicts accurately and can effectively extract and retain the knowledge useful for the target task; the other is that the compressed model has a smaller volume and faster inference speed. During the execution of the present invention, knowledge distillation based on Probes mainly provides a guide for searching for a small model applicable to the user intention recognition task. Compared with basic knowledge distillation (the small model only learns the output probability distribution of the large model), it can learn the knowledge in the intermediate layer of the original BERT model, thereby improving the performance; DARTS search can directly search for the neural network structure applicable to this task, reducing the complexity of manually designing the model or pruning and quantifying the original large model; further, the model obtained only by distillation and search has poor performance, so the optimal model obtained by search is fine-tuned to further improve the model performance.

[0086] An embodiment of the present invention further provides a user intention recognition method, which uses the obtained user intention recognition model to perform user intention recognition of whether it is a chat.

[0087] Figure 5 The main steps of the user intention recognition method in an embodiment are shown. Refer to Figure 5 As shown, the user intention recognition method includes: step S510, obtaining user input content; step S520, inputting the user input content into the user intention recognition model to obtain a user intention recognition result; step S530, according to the user intention recognition result, judging whether the user input content is a chat. If so, transfer to the intelligent customer service, otherwise transfer to the manual customer service.

[0088] The user intention recognition method of the present invention uses a user intention recognition model, which can accelerate the recognition of user intentions, timely judge user intentions, increase the self-service rate of the intelligent customer service system, and reduce the cost of manual customer service.

[0089] An embodiment of the present invention further provides a generating device for a user intention recognition model, which can be used to implement the method for generating a user intention recognition model described in any of the above embodiments. The features and principles of the generating method described in any of the above embodiments can be applied to the following generating device embodiments. In the following generating device embodiments, the features and principles regarding model generation that have been clarified will not be repeated.

[0090] Figure 6 Shows the main modules of the generating device of the user intention recognition model in an embodiment. Refer to Figure 6 As shown, the generating device 600 of the user intention recognition model includes: a knowledge distillation module 610, configured to train a teacher model including a multi-layer encoding network based on layer-by-layer knowledge distillation to obtain the output Logits of each layer of the encoding network of the target teacher model; a differentiable search module 620, configured to perform differentiable neural network architecture search based on a student model including a multi-layer convolutional network according to a target loss function including the cross-entropy loss between the output result of the student model and the true label of the training data, and the cross-entropy loss between the output results of each layer of the convolutional network and the output Logits of each layer of the encoding network, to obtain a target student model; a fine-tuning processing module 630, configured to fine-tune the target student model according to the cross-entropy loss between the output result of the target student model and the true label to obtain a user intention recognition model.

[0091] An embodiment of the present invention further provides a user intention recognition device, which can be deployed in the same device as the above-mentioned generating device of the user intention recognition model. The user intention recognition device can be used to recognize the user intention recognition method described in the above embodiment.

[0092] Figure 7 Shows the main modules of the user intention recognition device in an embodiment. Refer to Figure 7 As shown, the user intention recognition device 700 includes: a content acquisition module 710, configured to obtain user input content; an intention recognition module 720, configured to input the user input content into the user intention recognition model to obtain a user intention recognition result; a transfer processing module 730, configured to determine whether the user input content is a chat according to the user intention recognition result, and if so, transfer to an intelligent customer service, otherwise transfer to a human customer service.

[0093] The generating device / user intention recognition device of the user intention recognition model of the present invention can automatically obtain a target student model with reduced model parameter quantity and model size through layer-by-layer knowledge distillation and differentiable neural network architecture search, reducing the consumption of hardware resources; and perform fine-tuning after completing the model architecture search, significantly improving the model performance without retraining to obtain a user intention recognition model, realizing fully automatic intelligent acceleration of user intention recognition, increasing the self-service rate of the intelligent customer service system, and reducing the human customer service cost.

[0094] An embodiment of the present invention further provides an electronic device, including a processor and a memory, where an executable instruction is stored in the memory, and when the executable instruction is executed by the processor, it implements the generating method / user intention recognition method of the user intention recognition model described in any of the above embodiments.

[0095] The electronic device of the present invention can automatically obtain a target student model with reduced model parameter quantity and model size through layer-by-layer knowledge distillation and differentiable neural network architecture search, reducing the consumption of hardware resources; and perform fine-tuning after completing the model architecture search, significantly improving the model performance without the need for retraining, obtaining a user intention recognition model, realizing fully automatic intelligent acceleration for recognizing user intentions, increasing the self-service rate of the intelligent customer service system, and reducing the manual customer service cost.

[0096] Figure 8 It is a schematic structural diagram of the electronic device in the embodiment of the present invention. It should be understood that Figure 8 only various modules are schematically shown. These modules can be virtual software modules or actual hardware modules. The combination, splitting, and addition of other modules are all within the protection scope of the present invention.

[0097] As Figure 8 shown, the electronic device 800 is presented in the form of a general computing device. The components of the electronic device 800 include but are not limited to: at least one processing unit 810, at least one storage unit 820, a bus 830 connecting different platform components (including the storage unit 820 and the processing unit 810), a display unit 840, etc.

[0098] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 810, enabling the processing unit 810 to execute the steps of the method for generating the user intention recognition model / the method for recognizing user intentions described in any of the above embodiments. For example, the processing unit 810 can execute the steps as Figure 1 and Figure 5 shown.

[0099] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 8201 and / or a cache storage unit 8202, and may further include a read-only storage unit (ROM) 8203.

[0100] The storage unit 820 may further include a program / utilities 8204 having one or more program modules 8205. Such program modules 8205 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples.

[0101] The bus 830 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.

[0102] The electronic device 800 can also communicate with one or more external devices 900, which can be one or more of devices such as a keyboard, a pointing device, a Bluetooth device, etc. These external devices 900 enable a user to interact and communicate with the electronic device 800. The electronic device 800 can also communicate with one or more other computing devices, and the illustrated computer devices include a router and a modem. Such communication can be carried out through the input / output (I / O) interface 850. Also, the electronic device 800 can further communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 860. The network adapter 860 can communicate with other modules of the electronic device 800 through the bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.

[0103] An embodiment of the present invention also provides a computer-readable storage medium for storing a program, which when executed implements the method for generating a user intention recognition model / the user intention recognition method described in any of the above embodiments. In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the method for generating a user intention recognition model / the user intention recognition method described in any of the above embodiments.

[0104] When the storage medium of the present invention is executed again, it can automatically obtain a target student model with reduced model parameter quantity and model size through layer-by-layer knowledge distillation and differentiable neural network architecture search, reducing the consumption of hardware resources; and perform fine-tuning after completing the model architecture search, significantly improving the model performance without re-training, obtaining a user intention recognition model, realizing fully automatic intelligent acceleration for identifying user intentions, increasing the self-service rate of the intelligent customer service system, and reducing the cost of artificial customer service.

[0105] The program product can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto, and it can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.

[0106] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples of the readable storage medium include, but are not limited to: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0107] The readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0108] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device, for example, by connecting through the Internet using an Internet service provider.

[0109] The above content is a further detailed description of the present invention in conjunction with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for generating a user intention recognition model, characterized in that, it includes: Training a teacher model including a multi-layer encoding network based on layer-by-layer knowledge distillation to obtain the output Logits of each layer of the encoding network of the target teacher model; Based on a student model including a multi-layer convolutional network, performing differentiable neural network architecture search according to a target loss function including the cross-entropy loss between the output result of the student model and the true label of the training data, and the cross-entropy loss between the output results of each layer of the convolutional network and the output Logits of each layer of the encoding network, to obtain a target student model; Fine-tuning the target student model according to the cross-entropy loss between the output result of the target student model and the true label to obtain a user intention recognition model.

2. The generation method according to claim 1, characterized in that, The training of the teacher model including a multi-layer encoding network based on layer-by-layer knowledge distillation includes: Inserting a Probe classifier corresponding to each layer of the encoding network in the teacher model; Using the training data to train the teacher model based on knowledge distillation to obtain the output Logits of each layer of the encoding network through the Probe classifier.

3. The generation method according to claim 2, characterized in that, The obtaining of the output Logits of each layer of the encoding network of the target teacher model includes: Performing a preset number of trainings on the teacher model; Obtaining the optimal output of each layer of the encoding network in the preset number of trainings as the output Logits of each layer of the encoding network of the target teacher model.

4. The generation method according to claim 1, characterized in that, The teacher model adopts a Bidirectional Encoder Representations from Transformers (BERT) model; The candidate operators for each layer of the convolutional network include: Multiple convolutions with different convolution kernel sizes, multiple dilated convolutions with different convolution kernel sizes, average pooling, max pooling, Identity function, and Zero function.

5. The generation method according to claim 4, characterized in that, When performing differentiable neural network architecture search, each layer of the convolutional network is used as a search unit, and each search unit includes two input nodes, one output node, and multiple intermediate nodes; In each search unit, the two input nodes are the output nodes of the previous two search units, each intermediate node is connected to the output node, and each intermediate node has two incoming edges, and each incoming edge is selected from the candidate operators.

6. The generation method according to claim 1, characterized in that, The target loss function further includes an efficiency-aware loss for performing differentiable neural network architecture search; The formula of the target loss function is: wherein, is the cross-entropy loss between the output result of the student model and the true label, is the cross-entropy loss between the output results of each layer of the convolutional network and the output Logits of each layer of the encoding network, is the efficiency-aware loss, and γ and β are hyperparameters.

7. The generation method according to claim 6, characterized in that, The cross-entropy loss between the output result of the i-th convolutional network and the output Logits of the j-th encoding network is given by the formula: Among them, is the Probe classifier on the j-th layer encoding network, is the hidden representation of the j-th layer encoding network, is the output Logits of the j-th layer encoding network, is the Probe classifier on the i-th layer convolutional network, is the hidden representation of the i-th layer convolutional network, is the output result of the i-th layer convolutional network, and T is the temperature coefficient; where M is the number of samples of the training data, K is the number of convolutional network layers of the student model, and ω i,m is the normalized weight of the cross-entropy loss ; where y m is the label of the m-th sample, with the positive class being 1 and the negative class being 0.

8. The generation method according to claim 6, characterized in that: Among them, K is the number of convolutional network layers of the student model, and K max is the predefined maximum number of layers, and o i,j is the candidate operator to be searched, and α c is the network structure of the target student model, SIZE(.) is the parameter size, and FLOPs(.) is the number of floating-point operations for searching each candidate operator.

9. The generation method according to claim 6, characterized in that: where y m is the label of the m-th sample, with the positive class being 1 and the negative class being 0, and p(y m ) is the probability that the student model predicts the m-th sample as the positive class; When fine-tuning the target student model, p(y m ) is the probability that the target student model predicts the m-th sample as the positive class.

10. The generation method according to claim 1, characterized in that, The training data includes chitchat samples and non-chitchat samples; The user intention recognition model is used to identify whether the user input content is chitchat.

11. A method for user intention recognition, characterized in that, it includes: Obtain the user input content; Input the user input content into the user intention recognition model generated by the generation method described in any one of claims 1-10 to obtain the user intention recognition result; According to the user intention recognition result, judge whether the user input content is chitchat. If so, transfer to the intelligent customer service. If not, transfer to the human customer service.

12. A device for generating a user intention recognition model, characterized in that, it includes: A knowledge distillation module, which is used to train a teacher model including a multi-layer encoding network based on layer-by-layer knowledge distillation to obtain the output Logits of each layer of the encoding network of the target teacher model; A differentiable search module, which is used to perform differentiable neural network architecture search based on a student model including a multi-layer convolutional network according to a target loss function including the cross-entropy loss between the output result of the student model and the true label of the training data, and the cross-entropy loss between the output results of each layer of the convolutional network and the output Logits of each layer of the encoding network, to obtain a target student model; A fine-tuning processing module, which is used to fine-tune the target student model according to the cross-entropy loss between the output result of the target student model and the true label to obtain a user intention recognition model.

13. A user intention recognition device, characterized in that, it includes: A content acquisition module, which is used to obtain the user input content; An intention recognition module, which is used to input the user input content into the user intention recognition model generated by the generation method described in any one of claims 1-10 to obtain the user intention recognition result; A transfer processing module, which is used to judge whether the user input content is chitchat according to the user intention recognition result. If so, transfer to the intelligent customer service. If not, transfer to the human customer service.

14. An electronic device, characterized in that, it includes: A processor; A memory, in which executable instructions are stored; Wherein, when the executable instructions are executed by the processor, the generation method of the user intention recognition model described in any one of claims 1-10 is implemented, and / or, the user intention recognition method described in claim 11 is implemented.

15. A computer-readable storage medium for storing a program, characterized in that, when the program is executed by the processor, the generation method of the user intention recognition model described in any one of claims 1-10 is implemented, and / or, the user intention recognition method described in claim 11 is implemented.

Citation Information

Patent Citations

  • Pedestrian re-identification model construction method and device based on neural architecture search

    CN110852168A

  • Neural network structure search method and device, electronic equipment and storage medium

    CN111612134A