Method and apparatus for intent recognition in voice interaction

By constructing an intent recognition model, using document topic generation and BERT semantic representation to screen unique topics, combining multi-task learning and adversarial training, the accuracy and resource consumption problems of intent recognition in the speech interaction system are solved, and efficient recognition of unique intent is achieved.

CN114220426BActive Publication Date: 2025-07-18HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111370057.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-18
Publication Date
2025-07-18
Estimated Expiration
2041-11-18

AI Technical Summary

Technical Problem

Existing voice interaction systems cannot accurately identify user intentions, especially when facing dialects or complex contexts, and are highly consuming for multi-model management and resource consumption.

Method used

A intent recognition model is constructed, and the unique intent of text information is filtered through the document topic generation model and the isolated forest algorithm of BERT semantic representation. Combined with multi-task learning and adversarial training, the unique intent of text information is identified.

Benefits of technology

The unique intention recognition of voice information is realized, multi-model management and resource consumption are avoided, and the accuracy and efficiency of recognition are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114220426B_ABST
    Figure CN114220426B_ABST
Patent Text Reader

Abstract

The present application provides a method and apparatus for intent recognition in voice interaction. The method includes receiving voice information and converting the voice information into text information; inputting the obtained text information into an intent recognition model to obtain the unique intent information of the text information; the unique intent information is the intent information with the highest probability score among multiple intent information obtained after the intent recognition model recognizes the text information. The method of the present application can recognize the unique intent information of text information, thereby solving the problem that the existing voice interaction system still cannot accurately recognize the user's intent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to voice interaction technology, and in particular to a method and device for intent recognition in voice interaction. Background Art

[0002] With the continuous development of intelligent terminals, voice interaction systems have been widely applied in various intelligent terminals (such as smart TVs, in-vehicle navigation systems, smart speakers, etc.). In a voice interaction system, whether the user's expression can be well understood is related to the success or failure of the entire interaction process. Therefore, natural language understanding in a voice interaction system (understanding the text information converted from voice information, including intent recognition and slot extraction) is an important research direction of the voice interaction system.

[0003] Existing voice interaction systems generally complete intent recognition and understanding in natural language understanding by constructing a domain intent system. That is, through the constructed domain intent system for domain classification, and then performing intent recognition on the text information first, so as to analyze the intent information of the text information. For example, when a user says "I want to watch a movie", after the intelligent terminal converts the voice information into text information "I want to watch a movie", it first performs domain recognition (audio-visual playback domain), and then performs intent recognition on the text information (watch). However, the existing domain intent system is prone to identifying multiple different intents and domains for the text information. For example, if "I want to watch a movie" is spoken in a dialect, it may classify "I want to watch a movie" into the dialect interaction domain, and the intent may be recognized as voice interaction in the dialect.

[0004] Therefore, the existing voice interaction system still has the problem of being unable to accurately identify the user's intent. Summary of the Invention

[0005] The present application provides a method and device for intent recognition in voice interaction to solve the problem that the existing voice interaction system still cannot accurately identify the user's intent.

[0006] On the one hand, the present application provides a method for intent recognition in voice interaction, including:

[0007] Receiving voice information and converting the voice information into text information;

[0008] Inputting the converted text information into an intent recognition model to obtain the unique intent information of the text information;

[0009] The unique intent information is the intent information with the highest probability score among the multiple intent information obtained by the intent recognition model after recognizing the text information.

[0010] In one embodiment, it further includes:

[0011] Construct a training corpus for the initial intent recognition model, where each piece of text information in the training corpus has a unique theme and a unique action under the unique theme;

[0012] Train the initial intent recognition model with the text information in the training corpus of the initial intent recognition model to obtain the intent recognition model.

[0013] In one embodiment, the construction of the training corpus for the initial intent recognition model includes:

[0014] Use a document theme generation model to label each piece of text information in the text information library with N themes respectively, and output the probability score of each theme among the N themes, where N is equal to the number of themes identified in the text information, and N is an integer greater than or equal to 1;

[0015] Take K themes with probability scores of themes greater than a preset probability score as the finally labeled themes of each piece of text information, where K is an integer greater than or equal to 1 and K is less than or equal to N;

[0016] When K is greater than 1, use the isolation forest algorithm based on BERT semantic representation to detect whether the text information is an outlier of the first theme among the K themes;

[0017] When the text information is an outlier of the first theme, remove the annotation of the first theme from the K themes;

[0018] When the number of themes labeled for the text information after removing the annotation of the first theme is still greater than 1, use BERT similarity calculation to calculate the average similarity between the text information and non-outliers under each theme. When the average similarity between the text information and non-outliers under the second theme is the largest, determine the second theme as the unique theme of the text information;

[0019] Perform action partitioning on the text information according to the action words in the text information with a unique theme to obtain text information with a unique theme and a unique action under the unique theme;

[0020] Construct the training corpus of the initial intent recognition model with each piece of text information in the text information library having a unique theme and a unique action.

[0021] In one embodiment, it further includes:

[0022] Respond to the name definition operation to define the name of the unique theme of the text information;

[0023] The inputting the converted text information into the intent recognition model to obtain the unique intent information of the converted text information includes:

[0024] Input the converted text information into the intent recognition model to identify the unique name definition corresponding to the converted text information and the unique action under the unique name definition corresponding to the converted text information, so as to obtain the intent information of the text information.

[0025] In one embodiment, training the initial intent recognition model with the text information in the training corpus of the initial intent recognition model to obtain the intent recognition model includes:

[0026] Based on the multi-task learning method, train the initial intent recognition model with the text information in the training corpus of the initial intent recognition model to obtain the intent recognition model.

[0027] In one embodiment, the training the initial intent recognition model with the text information in the training corpus of the initial intent recognition model based on the multi-task learning method to obtain the intent recognition model includes:

[0028] Based on the initial intent recognition model, perform topic-unique self-attention semantic representation and topic-shared self-attention semantic representation on the text information corresponding to each topic in the training corpus respectively;

[0029] Based on the initial intent recognition model, splice the topic-unique self-attention semantic representation and the topic-shared self-attention semantic representation of all the text information in the training corpus, and perform recognition training on the intent information of each piece of text information in all the text information spliced with the topic-unique self-attention semantic representation and the topic-shared self-attention semantic representation to obtain the intent recognition model.

[0030] In one embodiment, it further includes:

[0031] Perform adversarial training on the initial intent recognition model during training.

[0032] On the other hand, the present application further provides an intent recognition device in voice interaction, including:

[0033] A voice processing module, configured to receive voice information and convert the voice information into text information;

[0034] An intent recognition module, configured to input the converted text information into the intent recognition model to obtain the unique intent information of the text information; the unique intent information is the intent information with the highest probability score among the multiple intent information obtained after the intent recognition model recognizes the text information.

[0035] On the other hand, the present application also provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0036] The memory stores computer-executable instructions;

[0037] The processor executes the computer-executable instructions stored in the memory to implement the intent recognition method in the voice interaction as described in the first aspect.

[0038] On the other hand, the present application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed, cause a computer to execute the intent recognition method in the voice interaction as described in the first aspect.

[0039] On the other hand, a computer program product of the present application includes a computer program, and when the computer program is executed by a processor, it implements the intent recognition method in the voice interaction as described in the first aspect.

[0040] The present application provides an intent recognition method and apparatus in voice interaction, which establish an intent recognition model and can recognize unique intent information for the text information converted from voice information. Different from the prior art in which multiple intent information of text information is easily recognized, the present application can accurately recognize the unique intent information of text information to achieve the purpose of accurately recognizing the user's intent. In addition, the present application only establishes one model for intent recognition of text information, which also avoids the problem of high costs in aspects such as management, deployment, and resources of multiple models. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.

[0042] Figure 1 It is a schematic diagram of an application scenario of the intent recognition method in voice interaction provided by the present application.

[0043] Figure 2 It is a schematic flowchart of the intent recognition method in voice interaction provided by an embodiment of the present application.

[0044] Figure 3 It is a partial schematic diagram of the construction of a training corpus in the intent recognition method in voice interaction provided by an embodiment of the present application.

[0045] Figure 4 It is a partial schematic diagram of the construction of a training corpus in the intent recognition method in voice interaction provided by an embodiment of the present application.

[0046] Figure 5Partial schematic diagram of constructing a training corpus in the intent recognition method in voice interaction provided by an embodiment of the present application.

[0047] Figure 6 Schematic diagram of an intent recognition device in voice interaction provided by an embodiment of the present application.

[0048] Figure 7 Schematic diagram of an electronic device provided by an embodiment of the present application.

[0049] Through the above-mentioned drawings, specific embodiments of the present disclosure have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present disclosure in any way, but to illustrate the concept of the present disclosure to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0050] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0051] Intelligent voice interaction is a new generation of interaction mode based on voice input, and users can obtain feedback results by speaking. With the continuous development of intelligent terminals, voice interaction systems have been widely applied in various intelligent terminals (such as smart TVs, in-vehicle navigation systems, smart speakers, etc.). In a voice interaction system, whether the user's expression can be well understood is related to the success or failure of the entire interaction process. Therefore, natural language understanding (understanding the text information obtained by converting voice information) in a voice interaction system is an important research direction of the voice interaction system.

[0052] When performing intent recognition and understanding, a text classification problem is usually modeled, and models such as CNN, LSTM, Transformer, and Bert are used to extract text features from text information, and then the intent recognition is completed through the Sigmoid function or the Softmax function. However, the original intent recognition method based on a single model can no longer cope with the ever-expanding business scenarios, so a domain intent system is constructed, which considers the intent recognition of text information as a task of two stages: domain classification and intent recognition. Domain classification is performed first, and then intent recognition is performed within the domain. However, the current domain intent system is constructed in a bottom-up manner, that is, the intent is constructed first, and then the constructed intent system is divided into corresponding domains. The domain intent system constructed in a bottom-up manner has many instabilities, and it is common to adjust a certain intent from domain A to domain B. For example, when a user says "I want to watch a movie", the intelligent terminal converts the voice information into the text information "I want to watch a movie", first performs domain recognition (audio and video playback domain), and then performs intent recognition (watch) on the text information. However, if "I want to watch a movie" is said in a dialect, then "I want to watch a movie" may be classified as the dialect interaction field, and the intent may be identified as voice interaction for audio and video playback in dialect.

[0053] Therefore, the existing domain intent system has a lot of instability and may not be able to accurately identify user intent. In addition, the existing domain intent system has a rule or model for intent recognition in each domain, so it will incur a considerable cost in terms of model management, deployment and resources.

[0054] Based on this, the present application provides a method and device for identifying intent in voice interaction, and establishes an intent recognition model, which can uniquely identify the subject and action of the text information converted from the voice information, and define the unique intent information of the text information by the unique subject and the unique action. Unlike the prior art that easily identifies multiple intent information of text information, the present application can accurately identify the unique intent information of the text information, so as to achieve the purpose of accurately identifying the user's intent. In addition, the present application only establishes one model for intent recognition of text information, and also avoids the high cost of multiple model management, deployment and resources.

[0055] The method for identifying intention in voice interaction provided in the present application is applied to electronic devices, such as computers, smart televisions, smart speakers, etc. Figure 1 This is a schematic diagram of the application of the method for intention recognition in voice interaction provided in the present application. In the figure, the electronic device receives voice information input by a user, converts the voice information into text information, and inputs the converted text information into an intention recognition model to obtain unique intention information of the converted text information.

[0056] Please refer to Figure 2 , an embodiment of the present application provides a method for intent recognition in voice interaction, including:

[0057] S210, receive voice information and convert the voice information into text information.

[0058] The voice information is input by the user to the electronic device (such as a smart TV). After receiving the voice information, the electronic device converts the voice information based on the support of its own software (such as voice recognition software) and hardware (built-in microphone) to obtain the corresponding text information. For example, when the user says "I want to watch a movie", the text information "I want to watch a movie" can be obtained after conversion.

[0059] S220, input the converted text information into an intent recognition model to obtain the unique intent information of the text information; the unique intent information is the intent information with the highest probability score among the multiple intent information obtained after the intent recognition model recognizes the text information.

[0060] A text information may have multiple themes, such as film and television, news, sports, finance, etc., but generally there is only one action, such as search, learning, watching / listening, etc. When the intent recognition model processes the text information, it determines the probability scores of the multiple intent information of the text information from the multiple themes of the text information and the action words in the text information, and takes the intent information with the highest probability score among the multiple intent information as the unique intent information of the text information.

[0061] For example, the text information is "I want to watch the B segment in A". A is a movie, and the B segment is a sports segment (such as a football segment) in the A movie. At this time, the text information has two themes of "film and television" and "sports", and there are two corresponding intents ("film and television search" and "sports search"). The role of the intent recognition model is to identify the intent information most relevant to the text information "I want to watch the B segment in A", that is, the unique intent information, which is the film and television search. After identifying the unique intent information of the text information, then execute the key parameter information under the unique intent based on the unique intent of the text information and the specific content of the text information (that is, the slot extraction part in natural language understanding), so as to complete the voice interaction process. Since the slot extraction part is not the focus of this application, it will not be elaborated in detail here.

[0062] Optionally, the intent recognition model is used to identify the unique theme and unique action of the text information, determine multiple intent information based on the unique theme and unique action of the text information, and then screen out the intent information with the highest probability score from the multiple intent information as the unique intent information of the text information. The unique theme can be understood as the theme most relevant to the content of the text information. For example, the unique theme of "I want to watch segment B in A" is "film and television", the unique action is "watch", and the determined unique intent information is "want to watch segment B in film and television A".

[0063] Before inputting the converted text information into the intent recognition model, it is also necessary to construct the intent recognition model. When constructing the intent recognition model, first obtain an initial intent recognition model, and then construct a training corpus for the initial intent recognition model. Use the text information in the training corpus of the initial intent recognition model to train the initial intent recognition model to obtain the intent recognition model. Among them, when constructing the training corpus of the initial intent recognition model, each piece of text information in the obtained training corpus has a unique theme and a unique action under the unique theme. The unique theme can be understood as the theme most relevant to the text information.

[0064] Specifically, when constructing the training corpus of the initial intent recognition model, a large amount of text information can be obtained from the existing domain intent system. These large amounts of text information are, for example, text information with simple themes such as "I want to watch a movie", "I want to watch stocks", "I want to watch football", as well as some text information with more and more complex themes. Then establish a text information library with the obtained text information. Then use the document topic generation model (Ida model) to label each piece of text information in the text information library with N themes, where N is equal to the number of themes identified in the text information, and N is an integer greater than or equal to 1. For example, the above-described "I want to watch segment B in A" is labeled with two themes (theme 1 and theme 2, and at this time, theme 1 and theme 2 have no specific name definitions).

[0065] After obtaining the N themes of the text information, use the document topic generation model to output the probability score of each theme among the N themes. Based on the probability scores of each theme among the N themes, screen out K themes whose theme probability scores are greater than the preset probability score from the N themes as the finally labeled themes of each piece of text information. K is an integer greater than or equal to 1 and less than or equal to N. For example, the preset probability score is 0.5, and among the N themes of the text information, the probability scores of two themes are greater than 0.5, then these two themes are the finally labeled themes of the text information. If there is only one theme finally labeled on the text information, then this one finally labeled theme is the unique theme of the text information, and there is no need to perform the following outlier detection steps to further screen the themes.

[0066] When there are multiple finally labeled topics in the text information, that is, when K > 1, the Isolation Forest algorithm based on BERT semantic representation is used to detect whether the text information is an outlier of the first topic (any one of the finally labeled K topics) among the multiple finally labeled topics. When the text information with multiple topics is an outlier of the first topic (the text information being an outlier of the first topic can be understood as the text information having a very low correlation with the first topic), the label of the first topic is removed from the multiple finally labeled topics (K topics). For example, if the text information "I want to watch the B segment in A" is an outlier of Topic 1, then the label of Topic 2 is removed from the two topics (Topic 1 and Topic 2) of the text information "I want to watch the B segment in A". Specifically, after each text information in the text information library is identified and labeled with N topics, there will be at least one text information under each of the N topics. For example, under Topic 1, there are "I want to watch the B segment in A", "I want to watch movie C", "I want to watch movie D", etc., and under Topic 2, there are "I want to watch the B segment in A", "I want to watch competition a", "I want to watch competition b", etc. The Isolation Forest algorithm based on BERT semantic representation can be used to detect whether "I want to watch the B segment in A" is an outlier of Topic 1 and Topic 2. If one or some topics among the multiple topics in the text information are removed based on the outlier, and the text information only corresponds to one topic, then the remaining single topic is defined as the unique topic of the text information, and there is no need to perform the step of using BERT similarity to determine the unique topic described below.

[0067] If one or some topics among the multiple finally labeled topics in the text information are removed based on the outlier, and the text information still corresponds to multiple topics, that is, when the number of labeled topics of the text information after removing the label of the first topic is still greater than 1, the mean similarity between the text information and the non-outliers under each topic is calculated using BERT similarity. When the mean similarity between the non-outliers under the second topic is the largest, the second topic is determined as the unique topic of the text information, and thus, the unique topic of the text information is determined.

[0068] A non-outlier is at least one other piece of text information under the theme. When calculating the average similarity between a piece of text information and non-outliers, the semantic representation vectors of each piece of text information under the theme are output using BERT encoding. Then, the cosine similarity is used to calculate at least one similarity between this piece of text information and the at least one other piece of text information. Finally, the average value of the sum of these at least one similarities is taken to obtain the average similarity between this piece of text information under the theme and non-outliers. For example, "I want to watch segment B in A" has two themes (Theme 1 and Theme 2). Under Theme 1, there are three pieces of text information in total: "I want to watch segment B in A", "I want to watch movie C", and "I want to watch movie D". After semantic representation based on BERT, the average similarity between "I want to watch segment B in A" and the other two pieces of text information reaches 0.9. Under Theme 2, there are three pieces of text information in total: "I want to watch segment B in A", "I want to watch game a", and "I want to watch game b". After semantic representation based on BERT, the average similarity between "I want to watch segment B in A" and the other two pieces of text information reaches 0.5. Then, the only theme of "I want to watch segment B in A" is Theme 1.

[0069] After determining the text information with a unique theme in the text information library, the staff can define the names of these unique themes. When defining the names, the electronic device responds to the name definition operation and defines the unique theme of the text information. The names that the unique theme can be defined as are, for example, film and television, news, sports, finance, etc. Then, according to the action words (such as watch, search, learn, etc.) in the text information with a unique theme, the actions of the text information are divided to obtain the text information with a unique theme and the unique action under this unique theme. Then, the training corpus of the initial intention recognition model is constructed with each piece of text information in the text information library that has a unique theme and a unique action. The framework for constructing the training corpus of the initial intention recognition model is, for example Figure 3 as shown, defining the name of the unique theme, such as Figure 3 shown, film and television, news, sports, and finance are four themes, and there are different actions under each theme. For example, under the "film and television" theme, there are actions such as "search", "learn", "watch / listen", etc.

[0070] Based on the name definition of the theme in the intention recognition model, step S220 can be correspondingly understood as: inputting the transformed text information into the intention recognition model to identify the unique name definition corresponding to the transformed text information, and identifying the unique action under the unique name definition corresponding to the transformed text information, so as to obtain the intention information of the text information.

[0071] After constructing the training corpus of the initial intent recognition model, the initial intent recognition model is trained with the text information in the training corpus of the initial intent recognition model. Optionally, during training, the initial intent recognition model is trained based on the multi-task learning method. That is, during the specific training of the initial intent recognition model, the intent recognition tasks of the text information belonging to different topics in the training corpus are generally used as different tasks to jointly learn and train. Specifically, during training, based on the initial intent recognition model, the text information corresponding to each topic in the training corpus is respectively subjected to private self-attention semantic representation (Private-Transformer) and shared self-attention semantic representation (Shared-Transformer). Based on the initial intent recognition model, the private self-attention semantic representation and the shared self-attention semantic representation of all the text information in the training corpus are concatenated, and intent information recognition training is performed on each piece of text information in all the text information concatenated with the private self-attention semantic representation and the shared self-attention semantic representation, so as to obtain the intent recognition model.

[0072] Please refer to Figure 4 , using Transformer as the encoder to perform self-attention representation on the embeddings of two pieces of text information ( Figure 4 For the sake of clarity, taking the self-attention representation and concatenation of two pieces of text information as an example, but it does not mean that two pieces of text information are processed each time in the solution). For each piece of text information in the training corpus (such as Figure 4 the shown Smple-M and Smple-N), first pass through the shared self-attention semantic representation (Shared-Transformer) and the private self-attention semantic representation (Private-Transformer) shown in Figure 4 , concatenate the two representation vectors obtained through Shared-Transformer and Private-Transformer, and pass through the fully connected layer and the Softmax layer to output the intent information recognition training of each piece of text information. The fully connected layer is used to recognize the classification of the intent information of the text information and the probability score of each type of intent information based on the representation vector after the concatenation of the text information, and obtain the intent information with the highest probability score as the unique intent information of the text information. The Softmax layer is used to process the output of the model.

[0073] Figure 4The shown Loss-m and Loss-n are the losses after intent recognition training. The initial intent recognition model is continuously optimized and updated according to the results of intent recognition training, and then intent recognition training is performed based on this training corpus until the loss after training no longer changes, completing the training of the initial intent recognition model and obtaining the intent recognition model.

[0074] Existing model training generally adopts the method of single-task learning, that is, only one task is learned at a time. For complex tasks, they will also be decomposed into simple and independent subtasks for separate learning, and then the learning results are combined. This single-task learning method ignores the characteristics of mutual correlation between tasks, so the effect of model training is poor. In this embodiment, when performing the initial intent recognition model, a multi-task learning (topic-exclusive self-attention representation plus topic-shared self-attention representation) method is adopted. During the process of model learning and training, the characteristics of mutual correlation between tasks are fully considered, making the learning and training effect of the initial intent recognition model better.

[0075] Optionally, when training the initial intent recognition model, adversarial training can also be performed on the initial intent recognition model during training to further modify the parameters of the initial intent recognition model according to the results of adversarial training. The method of adversarial training is, for example, adding a small perturbation to the input neurons of the Transformer (i.e., the output of the embedding), or adding adversarial training in the learning of the Shared-Transformer. The schematic diagram of adversarial training on the initial intent recognition model is as Figure 5 shown. r-m represents the perturbation, and this perturbation, for example, has multi-topic text information. Performing adversarial training on the initial intent recognition model can increase the robustness of the intent recognition model, making the intent recognition effect of the intent recognition model better.

[0076] Specifically, when performing adversarial training on the initial intent recognition model by adding a small perturbation to the input neurons of the Transformer (i.e., the output of the embedding), a perturbation r needs to be generated first. When generating this perturbation r, it is necessary to make the loss calculated under the parameters of the current model (the initial intent recognition model) the largest. The specific loss calculation methods are shown in Formulas 1 and 2. Formula 1: Formula 2: where, r em represents the perturbation value, ε is the step size, a hyperparameter with a small value, g represents the loss gradient, ‖g‖ represents the norm of the loss gradient, θ represents the current parameters of the intent recognition model, and p(y|x; θ) represents the probability that the model predicts the topic of sample x (text information) as y under the parameters θ.

[0077] Then, the loss of each sample after adding perturbations is calculated using Equation 3. Equation 3: where L adv (θ) represents the loss based on the perturbed samples, N is the total number of samples, r em,n represents the perturbation generated by Equation 1, s n + r em,n is the new perturbed sample, and p(y n |s n + r em,n ; θ) represents the probability that the model predicts the topic as y n + r em,n for the perturbed sample s n under the parameter θ.

[0078] When introducing an adversarial mechanism into the learning of Shared-Transformer for adversarial training of the model, it is necessary to first learn the Shared-Transformer feature network. Here, we adopt an interactive training method. First, a topic discriminator is given. The topic discriminator predicts the topic to which the sample belongs. Finally, the parameters of the Transformer feature network are updated to maximize the prediction loss of the topic discriminator. Among them, the prediction loss of the topic discriminator is the cross-entropy loss, which can be calculated according to Equation 4. Equation 4: where L adv-d (θ) represents the loss of the topic discriminator, N is the total number of samples,, s n is the sample under topic y n , and p(y n |s n ; θ) represents the probability that the topic predicted by the sample s n is y n under the discriminator parameter θ.

[0079] Then, based on the output of the Transformer feature network, the parameters of the topic discriminator are updated to minimize the prediction loss of the topic discriminator.

[0080] Finally, based on the loss calculated by Equation 5, the Shared-Transformer and Private-Transformer feature networks are updated to minimize the correlation loss between the two vectors output by the Shared-Transformer and Private-Transformer. Equation 5: where M is the number of topics, D m is the sample set of topic m, F s (x) is the encoding result of the shared-transformer, The encoding result of the private-transformer for theme m.

[0081] In summary, this embodiment provides a method for intent recognition in voice interaction. By constructing an intent recognition model, it recognizes the unique intent information of the text information converted from the voice information. Specifically, the intent recognition model recognizes the unique intent information of the text information converted from the voice information. Different from the prior art where multiple intent information of the text information is easily recognized, this application can accurately recognize the unique intent information of the text information to achieve the purpose of accurately recognizing the user's intent. In addition, this application only establishes one model for intent recognition of text information, which also avoids the problem of high costs in aspects such as management, deployment, and resources of multiple models.

[0082] Please refer to Figure 6 , an embodiment of this application also provides an intent recognition device 10 in voice interaction. The device 10 includes:

[0083] A voice processing module 11, configured to receive voice information and convert the voice information into text information.

[0084] An intent recognition module 12, configured to input the converted text information into the intent recognition model to obtain the unique intent information of the text information; the unique intent information is the intent information with the highest probability score among the multiple intent information obtained after the intent recognition model recognizes the text information.

[0085] The device 10 further includes:

[0086] A model construction module 13, configured to construct a training corpus for the initial intent recognition model. Each piece of text information in the training corpus has a unique theme and a unique action under the unique theme; the initial intent recognition model is trained with the text information in the training corpus of the initial intent recognition model to obtain the intent recognition model.

[0087] The model construction module 13 is specifically configured to use the document topic generation model to label each piece of text information in the text information library with N topics respectively, and output the probability score of each of the N topics. N is equal to the number of topics identified in the text information, and N is an integer greater than or equal to 1; the K topics with the probability score of the topic greater than the preset probability score are used as the finally labeled topics of each piece of text information. K is an integer greater than or equal to 1 and K is less than or equal to N; when K is greater than 1, the Isolation Forest algorithm based on BERT semantic representation is used to detect whether the text information is an outlier of the first topic among the K topics; when the text information is an outlier of the first topic, the annotation of the first topic is removed from the K topics; when the number of topics labeled for the text information after removing the annotation of the first topic is still greater than 1, the BERT similarity is used to calculate the average similarity between the text information and the non-outliers under each topic. When the average similarity between the text information and the non-outliers is the largest under the second topic, the second topic is determined as the unique topic of the text information; the text information is divided into actions according to the action words in the text information with the unique topic, and the text information with the unique topic and the unique action under the unique topic is obtained; the training corpus of the initial intent recognition model is constructed with each piece of text information in the text information library having the unique topic and the unique action.

[0088] The model construction module 13 is further configured to respond to the name definition operation and define the name of the unique topic of the text information. Correspondingly, the intent recognition module 12 is specifically configured to input the converted text information into the intent recognition model to identify the unique name definition corresponding to the converted text information, and identify the unique action under the unique name definition corresponding to the converted text information, so as to obtain the intent information of the text information.

[0089] The model construction module 13 is specifically configured to train the initial intent recognition model with the text information in the training corpus of the initial intent recognition model based on the multi-task learning method to obtain the intent recognition model. The model construction module 13 is specifically configured to perform topic-specific self-attention semantic representation and topic-shared self-attention semantic representation on the text information corresponding to each topic in the training corpus based on the initial intent recognition model; based on the initial intent recognition model, splice the topic-specific self-attention semantic representation and the topic-shared self-attention semantic representation of all the text information in the training corpus, and perform the recognition training of the intent information on each piece of text information in all the text information spliced with the topic-specific self-attention semantic representation and the topic-shared self-attention semantic representation to obtain the intent recognition model.

[0090] The model construction module 13 is further configured to perform adversarial training on the initial intent recognition model during training.

[0091] Please refer toFigure 7 In addition, the present application also provides an electronic device 20, including a processor 21 and a memory 22 communicatively connected to the processor 21. The memory 22 stores computer-executable instructions. The processor 21 executes the computer-executable instructions stored in the memory 22 to implement the intent recognition method in the voice interaction provided in any of the above embodiments.

[0092] The present application also provides a computer-readable storage medium storing computer-executable instructions, which when executed cause the computer-executable instructions to be executed by a processor to implement the intent recognition method in the voice interaction provided in any of the above embodiments.

[0093] The present application also provides a computer program product including a computer program, which when executed by a processor implements the intent recognition method in the voice interaction provided in any of the above embodiments.

[0094] It should be noted that the above computer-readable storage medium may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc. It may also be various electronic devices including one or any combination of the above memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, etc.

[0095] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including such element.

[0096] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0097] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0098] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general computer, a special computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0099] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0100] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0101] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.

Claims

1. A method for intent recognition in voice interaction, characterized in that, Including: Receiving voice information and converting the voice information into text information; Inputting the converted text information into an intent recognition model to obtain the unique intent information of the text information. The intent recognition model is based on the multi-task learning method, and the initial intent recognition model performs topic-exclusive self-attention semantic representation and topic-shared self-attention semantic representation on the text information corresponding to each topic in the training corpus of the initial intent recognition model respectively; Based on the topic-exclusive self-attention semantic representation and topic-shared self-attention semantic representation of all text information in the training corpus by the initial intent recognition model, splicing them, and performing recognition training on the intent information of each text information in all the text information spliced with the topic-exclusive self-attention semantic representation and topic-shared self-attention semantic representation. Each text information in the training corpus has a unique topic and a unique action under the unique topic; The unique intent information is the intent information with the highest probability score among the multiple intent information obtained after the intent recognition model recognizes the text information.

2. The method according to claim 1, characterized in that Also including: Constructing a training corpus for the initial intent recognition model.

3. The method according to claim 2, characterized in that, The construction of the training corpus for the initial intent recognition model includes: Using a document topic generation model to label each text information in the text information library with N topics respectively, and outputting the probability score of each of the N topics. N is equal to the number of topics recognized in the text information, and N is an integer greater than or equal to 1; Taking K topics with the probability score of the topic greater than the preset probability score as the finally labeled topics of each text information. K is an integer greater than or equal to 1 and less than or equal to N; When K is greater than 1, using the isolated forest algorithm based on BERT semantic representation to detect whether the text information is an outlier of the first topic among the K topics; When the text information is an outlier of the first topic, removing the annotation of the first topic from the K topics; When the number of topics labeled for the text information after removing the annotation of the first topic is still greater than 1, calculating the similarity mean between the text information and the non-outliers under each topic using BERT similarity. When the similarity mean between the text information and the non-outliers under the second topic is the largest, determining the second topic as the unique topic of the text information; Performing action division on the text information according to the action words in the text information with a unique topic to obtain text information with a unique topic and a unique action under the unique topic; Constructing the training corpus for the initial intent recognition model with each text information in the text information library having a unique topic and a unique action.

4. The method according to claim 3, wherein Also including: Responding to the name definition operation to define the name of the unique topic of the text information; The inputting the converted text information into the intent recognition model to obtain the unique intent information of the converted text information includes: Input the converted text information into the intent recognition model to identify the unique name definition corresponding to the converted text information and the unique action under the unique name definition corresponding to the converted text information, so as to obtain the intent information of the text information.

5. The method according to claim 1, wherein It further includes: Performing adversarial training on the initial intent recognition model during training.

6. An intent recognition device in voice interaction, characterized in that, It includes: A voice processing module, configured to receive voice information and convert the voice information into text information; An intent recognition module, configured to input the converted text information into the intent recognition model to obtain the unique intent information of the text information; The unique intent information is the intent information with the highest probability score among the multiple intent information obtained by the intent recognition model after recognizing the text information; The intent recognition model is based on the multi-task learning method, and the initial intent recognition model performs topic-exclusive self-attention semantic representation and topic-shared self-attention semantic representation on the text information corresponding to each topic in the training corpus of the initial intent recognition model respectively; Based on the topic-exclusive self-attention semantic representation and topic-shared self-attention semantic representation of all text information in the training corpus by the initial intent recognition model, splicing is performed, and intent information recognition training is performed on each text information in all text information spliced with the topic-exclusive self-attention semantic representation and topic-shared self-attention semantic representation; each text information in the training corpus has a unique topic and a unique action under the unique topic.

7. An electronic device, characterized in that, It includes: A processor and a memory communicatively connected to the processor; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory to implement the intent recognition method in the voice interaction as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the instructions are executed, the computer is caused to execute the intent recognition method in the voice interaction as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Intention recognition method in loud noise language environment

    CN109727598A

  • Intention recognition device, method and equipment based on hierarchical classification and storage medium

    CN111597320A