User intention recognition method, system, computing device and storage medium

By fusing single-intent and multi-intent recognition models, combining intent grouping relationships, and filtering and fusing intent prediction results, the problem of low intent recognition accuracy in multi-intent scenarios is solved, and higher intent recognition accuracy is achieved.

CN116189684BActive Publication Date: 2025-10-03TIANJIN AUTOHOME DATA INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310033745.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-10
Publication Date
2025-10-03
Estimated Expiration
2043-01-10

AI Technical Summary

Technical Problem

Existing intent recognition technology has low accuracy in multi-intent scenarios and cannot effectively handle the user's multiple expressions of intent, resulting in the problem of irrelevant content in telephone robots' voice interactions.

Method used

A method of fusing single-intent recognition model and multi-intent recognition model is adopted. A single intent is predicted by the first intent recognition model, and multiple intents are predicted by the second intent recognition model. Combined with the intent grouping relationship, the prediction results of the intent grouping are screened and fused to determine the end-user intent.

Benefits of technology

It improves the accuracy of intent recognition, solves the problem that different texts in multi-intent scenarios express the same result but have inconsistent true meanings, and improves the accuracy of intent recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189684B_ABST
    Figure CN116189684B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a method, system, computing device and storage medium for identifying user intentions. The method for identifying user intentions includes: obtaining voice data containing user answers and converting the voice data into text data; processing the text data through a first intention recognition model to predict a first intention indicating the user's intention; processing the text data through a second intention recognition model to predict a second intention indicating the user's intention, wherein the first intention and the second intention both belong to a user intention set, and the user intention set includes multiple intention groups; determining the intention corresponding to each intention group based on at least the relationship between the first intention and each intention group, and the second intention; fusing the intentions corresponding to each intention group to obtain the user's intention of the user's answer. According to this solution, the accuracy of intention recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a user intention recognition method, a computing device, and a storage medium. Background Art

[0002] In recent years, with the rapid development of artificial intelligence technology, intelligent outbound call robots have been increasingly used in corporate telemarketing, for example, in generating leads for car purchase intentions and in product marketing. Intent recognition is a core module in telebot systems, designed to automatically identify user responses and accurately understand user intent. Currently, intent recognition mostly focuses on a single intent within a sentence. However, in real-world scenarios, a user's sentence may have multiple meanings. Recognizing only a single intent can lead to issues such as irrelevant content and inability to grasp the user's core meaning during voice interaction. Therefore, effectively addressing the problem of multi-intent recognition has long been a key and challenging area of ​​intent recognition research.

[0003] Therefore, a new solution for identifying user intent is needed to improve the accuracy of intent recognition. Summary of the Invention

[0004] The present disclosure provides a solution for identifying user intention, in an effort to solve or at least alleviate at least one of the above problems.

[0005] According to one aspect of the present disclosure, a user intention recognition method is provided, including: obtaining voice data containing user answers and converting the voice data into text data; processing the text data through a first intention recognition model to predict a first intention indicating the user intention; processing the text data through a second intention recognition model to predict a second intention indicating the user intention, wherein the first intention and the second intention both belong to a user intention set, and the user intention set includes multiple intention groups; determining the intention corresponding to each intention group based on at least the relationship between the first intention and each intention group, and the second intention; and fusing the intentions corresponding to each intention group to obtain the user intention of the user answer.

[0006] Optionally, the method according to the present disclosure also includes: generating a user intention set based on the scenario of the user's answer; dividing the user intention set into multiple intention groups, wherein the union of the multiple intention groups is the user intention set, and the intention groups do not intersect with each other.

[0007] Optionally, in the method according to the present disclosure, the first intention predicted by the first intention recognition model is the intention with the largest probability value among all intentions; the second intention predicted by the second intention recognition model is multiple intentions with probability values ​​greater than a threshold among all intentions.

[0008] Optionally, in the method according to the present disclosure, the intention corresponding to each intention group is determined based at least on the relationship between the first intention and each intention group, and the second intention, including: if the first intention belongs to the intention group, then a first predetermined number of intentions belonging to the intention group are selected in sequence from the second intention as the intention corresponding to the intention group; if the first intention does not belong to the intention group, then a second predetermined number of intentions belonging to the intention group are selected in sequence from the second intention as the intention corresponding to the intention group.

[0009] Optionally, in the method according to the present disclosure, the intention corresponding to each intention group is determined based at least on the relationship between the first intention and each intention group, and the second intention, and also includes: selecting a first predetermined number or a second predetermined number of intentions belonging to the corresponding intention group from the second intention in order of the probability values ​​of each intention in the second intention from large to small.

[0010] Optionally, in the method according to the present disclosure, the intention corresponding to each intention group is determined based at least on the relationship between the first intention and each intention group, and the second intention, and also includes: setting a corresponding maximum number of intentions to be retained for each intention group, wherein the first predetermined number is the maximum number of intentions to be retained minus 1, and the second predetermined number is the maximum number of intentions to be retained.

[0011] Optionally, in the method according to the present disclosure, if the first intention belongs to the intention group, a first predetermined number of intentions belonging to the intention group are selected in sequence from the second intention as the intention corresponding to the intention group, including: if the first intention belongs to the intention group and the maximum number of intentions retained is 1, there is no need to select an intention from the second intention.

[0012] Optionally, in the method according to the present disclosure, the intents corresponding to the intent groups are fused to obtain the user intent of the user's answer, including: taking the union of the intents corresponding to the intent groups as the final user intent of the user's answer.

[0013] Optionally, the method according to the present disclosure also includes the step of training the first intent recognition model: obtaining the text data of the user's answer and marking the unique intent of each text data as a first training sample; inputting the text data into the initial first intent recognition model to predict the probability value belonging to each intent, and taking the intent with the largest probability value as the predicted first intent; using the marked data and predicted first intent in the first training sample to train the first intent recognition model until the training is completed to obtain a trained first intent recognition model.

[0014] Optionally, the method according to the present disclosure also includes the step of training a second intent recognition model: obtaining the text data of the user's answer and annotating the intent set of each text data, where the intent set contains multiple intents as a second training sample; inputting the text data into the initial second intent recognition model to predict the probability value belonging to each intent, and taking the intent with a probability value greater than a threshold as the predicted second intention; using the annotated data in the second training sample and the predicted second intention to train the second intent recognition model until the training is completed to obtain a trained second intent recognition model.

[0015] Optionally, in the method according to the present disclosure, when the scenario in which the user answers is a scenario for cleaning clues of the user's intention to buy a car, the intention grouping includes: an intention grouping indicating a tendency, an intention grouping for asking questions, and an intention grouping of a neutral category, wherein the intention grouping indicating a tendency includes at least the following intentions: intention to have already bought a car, intention to buy a car, affirmative intention, and negative intention; the intention grouping for asking questions includes at least the following intentions: intention to inquire about price, intention to inquire about identity; the intention grouping of a neutral category includes at least the following intentions: intention to hang up, intention to be temporarily inconvenient to communicate, and other intentions.

[0016] According to another aspect of the present disclosure, a user intention recognition system is provided, including: a text data acquisition unit, suitable for acquiring voice data containing user answers and converting the voice data into text data; an intention prediction unit, suitable for processing the text data through a first intention recognition model to predict a first intention indicating the user intention, and processing the text data through a second intention recognition model to predict a second intention indicating the user intention, wherein the first intention and the second intention both belong to a user intention set, and the user intention set includes multiple intention groups; an intention determination unit, suitable for determining the intention corresponding to each intention group based on at least the relationship between the first intention and each intention group, and the second intention; and further suitable for fusing the intentions corresponding to each intention group to obtain the user intention of the user answer.

[0017] According to another aspect of the present disclosure, a computing device is provided, comprising: one or more processor memories; one or more programs, wherein the one or more programs are stored in the memories and configured to be executed by one or more processors, and the one or more programs include instructions for executing any of the above methods.

[0018] According to yet another aspect of the present disclosure, a computer-readable storage medium storing one or more programs is provided. The one or more programs include instructions that, when executed by a computing device, cause the computing device to perform any of the methods described above.

[0019] In summary, according to the solution disclosed herein, the first intent recognition model and the second intent recognition model are integrated, the intent prediction results (the first intent of a single intent and the second intent of multiple intents) are grouped and integrated, and the second intent is screened by analyzing the relationship between the first intent and the grouped intent, thereby determining the final user intent. This can solve some problems that cannot be solved by existing multi-intent recognition models, such as situations where different texts have the same results recognized by the multi-intent model, but the real meanings are actually different. The solution disclosed herein can effectively improve the accuracy of intent recognition.

[0020] The above description is only an overview of the technical solution of the present disclosure. In order to more clearly understand the technical means of the present disclosure, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present disclosure more obvious and easy to understand, the specific implementation methods of the present disclosure are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] To achieve the above and related purposes, certain illustrative aspects are described herein in conjunction with the following description and accompanying drawings, which indicate various ways in which the principles disclosed herein may be practiced, and all aspects and their equivalents are intended to fall within the scope of the claimed subject matter. The above and other objects, features, and advantages of the present disclosure will become more apparent by reading the following detailed description in conjunction with the accompanying drawings. Throughout this disclosure, the same reference numerals generally refer to the same parts or elements.

[0022] Figure 1 shows a schematic diagram of a user intent recognition system 100 according to some embodiments of the present disclosure;

[0023] Figure 2 shows a schematic diagram of a computing device 200 according to some embodiments of the present disclosure;

[0024] Figure 3 A flow chart of a method 300 for identifying user intent according to some embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0025] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0026] As described in the background technology, single intent recognition itself is a text classification problem. Common methods include traditional machine learning methods based on logistic regression and naive Bayes, as well as text classification methods based on deep learning that have been commonly used in recent years, such as convolutional neural networks (CNN), recurrent neural networks (RNN), and BERT models. A simple and feasible method is to segment the user input text by punctuation marks to obtain several clauses divided according to the punctuation marks, and then apply a single intent recognition model to each clause, and treat the set of single intents obtained from all clauses as the result of multiple intents. However, this method relies on the correctness of punctuation marks. In the telephone robot scenario, the punctuation marks in the speech recognition (Automatic Speech Recognition, ASR) text recorded by the user are often incorrect, which affects the accuracy of the method. On the other hand, this method cannot handle the situation where a clause has multiple meanings.

[0027] One approach to addressing these shortcomings is to directly employ a multi-intent recognition model. This is typically accomplished by transforming the multi-intent recognition problem into a multi-label text classification problem, where the number of labels equals the total number of intent categories, and each label has two classifications (yes or no). This multi-label text classification model can then identify multiple intents corresponding to a single input text. However, this approach can only derive a few intents from the user's input text and, in some cases, cannot effectively capture the user's true intent. For example, in a car buying and selling scenario, consider the following two sentences: the first sentence is "I've already bought a car, and I'm still looking for another one." The second sentence is "I've already bought a car, but I want to buy another one." The multi-intent recognition model yields the same results for both sentences: "Intention to have already bought a car" and "Intention to buy a car." However, in reality, the user's true intention in the first sentence is "Intention to have already bought a car," and the second sentence is "Intention to buy a car."

[0028] To address the above issues, the present disclosure proposes an intent recognition solution that integrates a single-intent model and a multi-intent model, and provides a method for fusing the intent results obtained by these two intent recognition models to improve the accuracy of multi-intent recognition.

[0029] Figure 1 FIG. 1 shows a schematic diagram of a user intention recognition system 100 according to the present disclosure. Figure 1 As shown, the user intention recognition system 100 includes: a text data acquisition unit 110 , an intention prediction unit 120 and an intention determination unit 130 .

[0030] The text data acquisition unit 110 acquires voice data containing the user's response and converts the voice data into text data. In some embodiments, the voice data is transcribed into text data containing the user's response using automated speech recognition (ASR). The text data acquisition unit 110 then forwards the text data to the coupled intent prediction unit 120.

[0031] The intention prediction unit 120 performs intention recognition using a single-intention model and a multi-intention model, respectively, to predict multiple intentions for the text data. It should be understood that the term "multiple" in this disclosure generally refers to "two" or "more than two."

[0032] Specifically, the intent prediction unit 120 processes the text data using a first intent recognition model to predict a first intent indicating the user's intent, and also processes the text data using a second intent recognition model to predict a second intent indicating the user's intent. The first intent and the second intent both belong to a user intent set, and the first intent is a single intent, while the second intent includes multiple intents, i.e., the second intent is an intent set.

[0033] In addition, the user intention set includes all possible user intentions that the user answers in the current question-answering scenario. According to some embodiments, the user intention set also includes multiple intention groups. For example, the user intention set is recorded as A t , group it and divide it into n groups A t1 ,A t2 ,…A tn , the union of n groups is equal to the user intention set, but the n sets do not intersect with each other, denoted as A t =A t1 ∪A t2 ∪…∪A tn ,

[0034] Afterwards, the intention prediction unit 120 sends the first intention and the second intention to the intention determination unit 130 coupled thereto, which determines the intention corresponding to each intention group based at least on the relationship between the first intention and each intention group, and the second intention, and then integrates the intentions corresponding to each intention group to finally obtain the user intention answered by the user.

[0035] According to the user intention recognition system 100 disclosed in the present invention, the first intention recognition model and the second intention recognition model are integrated, and the obtained intent prediction results include the first intention of a single intention and the second intention of multiple intentions. Afterwards, all possible intentions of the user are grouped, and the intent prediction results are grouped and fused. According to the relationship between the first intention and the intention grouping, the multiple intentions in the second intention are screened, and then the final user intention is obtained by fusion. It can solve some problems that cannot be solved by using the previous multi-intent recognition model, such as the situation where different texts have the same results recognized by the multi-intent model, but the real meaning is actually different. The use of system 100 can improve the accuracy of intent recognition.

[0036] The user intention recognition system 100 according to the present disclosure can be applied to various voice interaction scenarios, for example, it can be deployed in a telephone robot system. After obtaining the user's answer, the telephone robot system will hand it over to the user intention system 100 to determine the intention corresponding to the user's answer. After that, the telephone robot system can perform subsequent operations such as process jumps based on the recognized intention. Of course, the user intention recognition system 100 can also be deployed in smart home devices to determine the user's true intention by analyzing the user's voice answer so as to make a correct response. This disclosure does not impose too many restrictions on this.

[0037] According to the present disclosure, the user intent recognition system 100 may be displayed via one or more computing devices.

[0038] Figure 2 FIG. 2 shows a structural block diagram of a computing device 200 according to an embodiment of the present disclosure.

[0039] like Figure 2 As shown, in a basic configuration 202, computing device 200 typically includes system memory 206 and one or more processors 204. A memory bus 208 may be used for communication between processor 204 and system memory 206.

[0040] Depending on the desired configuration, the processor 204 can be any type of processor, including, but not limited to, a microprocessor (μP), a microcontroller (μC), a digital signal processing unit (DSP), or any combination thereof. The processor 204 can include one or more levels of cache, such as a level 1 cache 210 and a level 2 cache 212, a processor core 214, and registers 216. An example processor core 214 can include an arithmetic logic unit (ALU), a floating point unit (FPU), a digital signal processing (DSP) core, or any combination thereof. An example memory controller 218 can be used with the processor 204, or in some implementations, the memory controller 218 can be an internal part of the processor 204.

[0041] Depending on the desired configuration, system memory 206 can be any type of memory, including but not limited to volatile memory (such as RAM), non-volatile memory (such as ROM, flash memory, etc.), or any combination thereof. Physical memory in a computing device generally refers to volatile RAM, and data from a disk must be loaded into physical memory before it can be read by processor 204. System memory 206 can include an operating system 220, one or more applications 222, and program data 224. In some embodiments, application 222 can be arranged so that one or more processors 204 execute instructions on the operating system using program data 224. Operating system 220 can be, for example, Linux, Windows, etc., and includes program instructions for handling basic system services and performing hardware-dependent tasks. Application 222 includes program instructions for implementing various user-desired functions. Application 222 can be, for example, a browser, instant messaging software, software development tools (such as an integrated development environment IDE, a compiler, etc.), etc., but is not limited thereto.

[0042] When computing device 200 is started, processor 204 reads and executes program instructions from operating system 220 from memory 206. Applications 222 run on operating system 220, utilizing interfaces provided by operating system 220 and the underlying hardware to implement various user-desired functions. When a user launches application 222, it is loaded into memory 206, and processor 204 reads and executes the program instructions from memory 206.

[0043] The computing device 200 also includes a storage device 232, which includes a removable storage 236 (such as a CD, DVD, USB flash drive, removable hard disk, etc.) and a non-removable storage 238 (such as a hard disk drive HDD, etc.). The removable storage 236 and the non-removable storage 238 are both connected to the storage interface bus 234.

[0044] The computing device 200 may also include a storage interface bus 234. The storage interface bus 234 enables communication from storage devices 232 (e.g., removable storage 236 and non-removable storage 238) via the bus / interface controller 230 to the basic configuration 202. At least a portion of the operating system 220, applications 222, and program data 224 may be stored on the removable storage 236 and / or the non-removable storage 238 and loaded into the system memory 206 via the storage interface bus 234 when the computing device 200 is powered on or when the application 222 is to be executed, and executed by the one or more processors 204.

[0045] The computing device 200 may also include an interface bus 240 that facilitates communication from various interface devices (e.g., output devices 242, peripheral interfaces 244, and communication devices 246) to the basic configuration 202 via the bus / interface controller 230. Example output devices 242 include a graphics processing unit 248 and an audio processing unit 250. These can be configured to facilitate communication with various external devices such as a display or speakers via one or more A / V ports 252. Example peripheral interfaces 244 may include a serial interface controller 254 and a parallel interface controller 256, which can be configured to facilitate communication with external devices such as input devices (e.g., a keyboard, mouse, pen, voice input device, touch input device) or other peripherals (e.g., a printer, scanner, etc.) via one or more I / O ports 258. Example communication devices 246 may include a network controller 260, which can be arranged to facilitate communication with one or more other computing devices 262 via a network communication link via one or more communication ports 264.

[0046] A network communication link can be an example of a communication medium. Communication media can generally be embodied as computer-readable instructions, data structures, program modules in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium. A "modulated data signal" can be a signal in which one or more of a data set or a change therein can be carried out in a manner that encodes information in the signal. As non-limiting examples, communication media can include wired media such as a wired network or a dedicated line network, and various wireless media such as sound, radio frequency (RF), microwave, infrared (IR) or other wireless media. The term computer-readable medium as used herein can include both storage media and communication media.

[0047] The computing device 200 can be implemented as a personal computer including desktop and notebook computer configurations. Of course, the computing device 200 can also be implemented as part of a small-sized portable (or mobile) electronic device, such as a cellular phone, a digital camera, a personal digital assistant (PDA), a personal media player device, a wireless network browsing device, a personal head-mounted device, an application-specific device, or a hybrid device that can include any of the above functions. It can even be implemented as a server, such as a file server, a database server, an application server, and a web server. The embodiments of the present disclosure are not limited to this.

[0048] In an embodiment of the present disclosure, the computing device 200 is configured to execute the user intent recognition method 300 according to the present disclosure. The application 222 deployed on the operating system includes multiple program instructions for executing one or more of the above methods. These program instructions can instruct the processor 204 to execute the above methods of the present disclosure to accurately extract the user intent from the user's voice response.

[0049] Figure 3 1 shows a flow chart of a method 300 for identifying user intent according to some embodiments of the present disclosure. In some embodiments, the method 300 is performed using the aforementioned system 100. It should be understood that the descriptions of the method 300 and the system 100 complement each other, and the relevant parts will not be repeated here.

[0050] Furthermore, this solution is described in detail using the example of cleaning user car purchase intention leads during intelligent outbound calls. It should be understood that this is merely an example, and those skilled in the art can apply it to other telephony robot scenarios or other voice interaction scenarios, all of which are within the scope of this disclosure.

[0051] like Figure 3 As shown, the method 300 begins at step S310 .

[0052] In step S310 , voice data containing the user's answer is acquired and converted into text data.

[0053] According to some embodiments, the user's voice data is transcribed into text using automated speech recognition (ASR). The speech recognition model may be implemented by calling a third-party interface or other methods, which are not limited by this disclosure. For example, if the audio content of a certain user voice data is "I want to buy a car", the converted user response text will be: "I want to buy a car."

[0054] In addition, according to some embodiments, method 300 further includes: generating a user intention set based on the scenario in which the user answers. The set of all possible intentions answered by the user in the scenario is taken as the user intention set, denoted as A t , the number of intentions is C, then any intention A∈A t ={A1,A2,...,A C In the current scenario, for example, a user intent set can be generated as {"intention to have purchased a car", "intention to purchase a car", "affirmative intention", "negative intention", "inquiry about price", "inquiry about identity", "hang up intention", "temporarily inconvenient to communicate", "other intentions"}, of course not limited to this.

[0055] Then, the user intention set is divided into multiple intention groups, where the union of multiple intention groups is the user intention set, and each intention group does not intersect with each other. t Divide into n groups, denoted as A t1 ,A t2 ,…A tn , A t =A t1 ∪A t2 ∪…∪A tn ,and Continuing with the previous example, all intents are divided into three categories: intent indicating preference, intent asking questions, and neutral intent. The intent indicating preference includes at least the following intents: already bought a car, intending to buy a car, affirmative intent, and negative intent; the intent asking questions includes at least the following intents: asking about price and asking about identity; and the neutral intent includes at least the following intents: hanging up, temporarily unavailable, and other. This disclosure does not restrict the specific method used to group user intent sets, and can be configured based on specific scenarios.

[0056] In addition, in some other embodiments, a maximum number of reserved intents corresponding to each intent group is set, which is recorded as n. mi , i=1,2,…,n.

[0057] In step S320 , the text data is processed by a first intent recognition model to predict a first intent indicating the user's intent.

[0058] The first intent recognition model generally adopts a single-label multi-classification text classification algorithm. After feature extraction and classification of the input text data, it obtains the predicted probability value of each intent (that is, the probability of predicting that the current user intention belongs to any intention A), and the sum of the predicted probability values ​​of each intention is 1. The intention with the largest probability value is taken as the first intention predicted. In other words, the first intent predicted by the first intent recognition model is the one with the largest probability value among all the intentions in the user intention set, recorded as the first intention A. sp ∈A t Assume that the user intent set is {“already bought a car intention”, “will buy a car intention”, “affirmed intention”}. For the text “I bought a car, and I want to buy another car.”, the first intent recognition model predicts the following probabilities for all intents: {“already bought a car intention”: 0.7, “will buy a car intention”: 0.29, “affirmed intention”: 0.01}, so the first intent is determined to be “already bought a car intention”.

[0059] According to the present disclosure, method 300 further includes the step of training a first intent recognition model, specifically including the following three steps.

[0060] The first step is to obtain the text data of the user's answer and mark the unique intention of each text data as the first training sample.

[0061] We obtain multiple voice data from the phone recordings of the outbound call robot and the user, and convert them into text data using ASR technology. We perform single intent annotation on each text data, analyze the main intent in the text, and annotate it to obtain a unique intent A. st ∈A t For example, for a piece of text data, "I bought a car, and I'm going to buy another car.", its labeled data is "intent to buy a car." The text data and its corresponding labeled data are used as the first training sample to train the first intent recognition model.

[0062] In the second step, the text data is input into the initial first intent recognition model to predict the probability of belonging to each intent, and the intent with the highest probability value is used as the predicted first intent. According to some embodiments of the present disclosure, the first intent recognition model may include a feature extraction module and a classification prediction module. The feature extraction module may adopt a CNN, RNN, or BERT model, and the classification prediction module may adopt a Softmax processing layer, etc., which are not limited by the present disclosure.

[0063] In the third step, the first intent recognition model is trained using the labeled data and the predicted first intent in the first training sample until the training is completed to obtain a trained first intent recognition model. According to some embodiments, a loss function is calculated based on the labeled intent and the predicted first intent. In this case, the loss function is as follows:

[0064]

[0065] Among them, y j and p j They respectively represent the true label value of whether the text data belongs to the j-th intent (the value is 1 if it does, and the value is 0 if it does not), and the probability value of the text data belonging to the j-th intent predicted by the first intent recognition model. C is the number of intents in the intent set.

[0066] In step S330 , the text data is processed by a second intent recognition model to predict a second intent indicating the user's intent.

[0067] The second intent recognition model generally adopts a multi-label and multi-classification text classification algorithm. After feature extraction and classification of the input text data, it obtains the predicted probability value of each intent (that is, the probability of predicting that the current user intention belongs to any intention A), and takes the set of intents with probability values ​​greater than a threshold (generally 0.5) as the predicted second intent, denoted as Furthermore, in the second intent recognition model, the sum of the predicted probabilities for each intent does not need to be 1. In other words, the second intent predicted by the second intent recognition model is the set of intents whose probabilities exceed a threshold. Assume that the user intent set is {"already bought a car intention", "will buy a car intention", "affirmative intention"}. For the text data "I bought a car, and I'm still buying a car.", the second intent recognition model predicts the following probabilities for all intents: {"already bought a car intention": 0.9, "will buy a car intention": 0.9, "affirmative intention": 0.05}. Therefore, the second intent is determined to be {"already bought a car intention", "will buy a car intention"}.

[0068] According to the present disclosure, method 300 further includes the step of training a second intent recognition model, specifically including the following three steps.

[0069] In the first step, the text data of the user's answer is obtained and the intent set of each text data is marked. The intent set contains multiple intents as the second training sample.

[0070] We obtain multiple voice data from the phone recordings between the outbound call robot and the user, and convert them into text data using ASR technology. We perform multi-intent annotation on each text data, analyze all the intentions of the literal meaning in the text data, and annotate them to obtain an intent set. For example, for the text data "I bought a car, and I'm going to buy a car again.", its annotated data is {"intention to have bought a car", "intention to buy a car"}. The text data and its corresponding annotated data are used as the second training sample to train the second intent recognition model.

[0071] In the second step, the text data is input into the initial second intent recognition model to predict the probability value belonging to each intent, and the intent with a probability value greater than a threshold (generally 0.5) is used as the predicted second intent. According to some embodiments of the present disclosure, the second intent recognition model may include a feature extraction module and a classification prediction module, wherein the feature extraction module may adopt CNN, RNN, Bert model, and the classification prediction module may adopt Softmax processing layer, etc., which is not limited by the present disclosure. The classification prediction module predicts the probability that the user intent of the input text data belongs to any intent A, and takes the set of intents with probability values ​​greater than the threshold as the predicted second intent.

[0072] In the third step, the second intent recognition model is trained using the labeled data and the predicted second intent in the second training sample until the training is completed and a trained second intent recognition model is obtained. According to some embodiments, a loss function is calculated based on the labeled intent set and the predicted second intent. In this case, the loss function is as follows:

[0073]

[0074] Among them, y j and p j They respectively represent the true label value of whether the text data belongs to the j-th category of intent (if it does, the value is 1, if it does not, the value is 0), and the probability value of the text data belonging to the j-th category of intent predicted by the second intent recognition model.

[0075] Two sets of intentions A are obtained through S320 and S330 sp and A mp How to combine these two sets of intents to obtain the final intent has a significant impact on the accuracy of multi-intent recognition. A simple approach is to take the union of the two sets of intents, but as shown in the previous analysis, this approach does not solve the problem of inaccurate intent recognition in text data such as "I bought a car, and I'm still buying a car."

[0076] Therefore, in step S340 , the intent corresponding to each intent group is determined based on at least the relationship between the first intent and each intent group, and the second intent.

[0077] First, traverse all intent groups A ti , i=1,2,…,n, judge the first intention A sp Whether it belongs to intent group A ti Afterwards, according to the first intention A sp Group A with each intention ti The relationship between the intention groups A and ti Regarding the corresponding intents, it should be noted that each intent group may correspond to one or more intents, or may not have any intents (ie, the intent corresponding to the intent group is empty).

[0078] Specifically, the relationship between the first intent and a certain intent group includes: the first intent belongs to the intent group or the first intent does not belong to the intent group. The following further describes these two relationships.

[0079] (1) If the first intention A sp Belongs to the intent group A ti , then from the second intention A mp The first predetermined number of intentions belonging to the intention group are selected in sequence as the intention group A ti According to some embodiments, "selecting in sequence" is defined as selecting a first predetermined number of intents belonging to the corresponding intent group from the second intent in descending order of the probability values ​​of the intents in the second intent. In addition, the first predetermined number is the maximum number of intents to be retained minus 1, that is, n mi -1. In other words, if A sp Belongs to intent group A ts , denoted as i=s, then the second intention A is retained mp Ats The first n mi -1 intent, and the intent retention order is in descending order of the probability values ​​of each intent output by the second intent recognition model.

[0080] In particular, when the intent group A ts When the maximum number of corresponding intents is 1, there is no need to start from the second intent A mp Select the intention.

[0081] Afterwards, a first predetermined number of intents selected from the second intent are merged with the first intent as the final corresponding intent of the intent group.

[0082] (2) If the first intention A sp Does not belong to the intention group A ti , then from the second intention A mp The second predetermined number of intentions belonging to the intention group are selected in sequence as the intention group A ti According to some embodiments, “selecting in sequence” is defined as selecting a second predetermined number of intents belonging to the corresponding intent group from the second intent in descending order of the probability values ​​of the intents in the second intent. In addition, the second predetermined number is the maximum number of intents to be retained, i.e., n mi In other words, if i≠s, then keep the second intention A mp A ts The first n of the set grouping mi Intents are retained in the same order as above.

[0083] In step S350, the intents corresponding to the intent groups are fused to obtain the user intent answered by the user.

[0084] In some embodiments, the union of the intents corresponding to the intent groups is taken as the final user intent of the user's answer.

[0085] Continuing with the previous example, in this embodiment, the user intention set A t= A1∪A2∪A3. Intent group A1 represents the tendency category, which includes {"already bought a car intention", "will buy a car intention", "affirmative intention", "negative intention", and the maximum number of retained intents is 1. Intent group A2 represents the question category, which includes {"inquiry about price", "inquiry about identity", and the maximum number of retained intents is 2. Intent group A3 represents the neutral other category, which includes {"hang up intention", "temporarily unavailable for communication", "other intention", and the maximum number of retained intents is 1. For the input text data "I bought a car, and I'm still buying a car.", the results obtained by the first intent recognition model are: {"already bought a car intention"}, and the results obtained by the second intent recognition model are: {"already bought a car intention", "will buy a car intention"}. After grouping and fusion, the final user intent determined is {"already bought a car intention"}.

[0086] Furthermore, several examples of text data are given, including their first intention, second intention, and the user intention finally determined after grouping and fusion, as shown in Table 1.

[0087] Table 1 Intent recognition examples

[0088]

[0089] According to the method 300 disclosed in the present invention, the first intent recognition model and the second intent recognition model are integrated, and the obtained intent prediction results include the first intention of a single intent and the second intention of multiple intents. Afterwards, all possible intentions of the user are grouped, and the intent prediction results are grouped and integrated. According to the relationship between the first intention and the intention grouping, the multiple intentions in the second intention are screened, and then the final user intention is obtained by integration. This solution can solve some problems that the existing multi-intent recognition model cannot solve, such as the situation where different texts have the same results recognized by the multi-intent model, but the real meanings are actually different. According to the method 300 disclosed in the present invention, the accuracy of intent recognition can be improved.

[0090] This disclosure also discloses:

[0091] A8. A method as described in any one of A1-7, wherein the fusing of the intents corresponding to the intent groups to obtain the user intent of the user's answer includes: taking the union of the intents corresponding to the intent groups as the final user intent of the user's answer.

[0092] A9. The method as described in any one of A1-8 also includes the step of training the first intent recognition model: obtaining the text data of the user's answer and marking the unique intention of each text data as the first training sample; inputting the text data into the initial first intent recognition model to predict the probability value of each intent, and taking the intent with the largest probability value as the predicted first intention; using the marked data in the first training sample and the predicted first intention to train the first intent recognition model until the training is completed to obtain the trained first intent recognition model.

[0093] A10. The method as described in any one of A1-9 further includes the step of training the second intent recognition model: obtaining the text data of the user's answer and annotating the intent set of each text data, wherein the intent set contains multiple intents as a second training sample; inputting the text data into the initial second intent recognition model to predict the probability value belonging to each intent, and taking the intent with a probability value greater than a threshold as the predicted second intention; using the annotated data in the second training sample and the predicted second intention to train the second intent recognition model until the training is completed to obtain a trained second intent recognition model.

[0094] A11. The method as described in A2, wherein, when the scenario in which the user answers is a scenario for cleaning clues of the user's intention to buy a car, the intention grouping includes: an intention grouping indicating a tendency, an intention grouping for asking questions, and an intention grouping of a neutral category, wherein the intention grouping indicating a tendency includes at least the following intentions: intention to have already bought a car, intention to buy a car, affirmative intention, and negative intention; the intention grouping for asking questions includes at least the following intentions: intention to inquire about price, intention to inquire about identity; the intention grouping of the neutral category includes at least the following intentions: intention to hang up, intention to be temporarily inconvenient to communicate, and other intentions.

[0095] The various techniques described herein may be implemented in conjunction with hardware or software, or a combination thereof. Thus, the methods and apparatus of the present disclosure, or certain aspects or portions of the methods and apparatus of the present disclosure, may take the form of program code (i.e., instructions) embedded in a tangible medium, such as a removable hard disk, a USB flash drive, a floppy disk, a CD-ROM, or any other machine-readable storage medium, wherein when the program is loaded into a machine such as a computer and executed by the machine, the machine becomes an apparatus for practicing the present disclosure.

[0096] When program code is executed on a programmable computer, the computing device generally includes a processor, a storage medium readable by the processor (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. The memory is configured to store the program code; the processor is configured to execute the user intent identification method of the present disclosure according to the instructions in the program code stored in the memory.

[0097] By way of example and not limitation, readable media include readable storage media and communication media. Readable storage media store information such as computer-readable instructions, data structures, program modules, or other data. Communication media typically embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and include any information delivery medium. Combinations of any of the above are also included within the scope of readable media.

[0098] In the description provided herein, the algorithms and displays are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the examples of the present disclosure. Based on the above description, it is apparent that the structure required for constructing such systems is suitable. In addition, the present disclosure is not directed to any specific programming language. It should be understood that various programming languages ​​can be utilized to implement the content of the present disclosure described herein, and the above description of specific languages ​​is intended to disclose preferred embodiments of the present disclosure.

[0099] In the description provided herein, numerous specific details are described. However, it is understood that embodiments of the present disclosure may be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.

[0100] Similarly, it should be understood that in order to streamline the present disclosure and aid in understanding one or more of the various disclosed aspects, in the above description of exemplary embodiments of the present disclosure, various features of the present disclosure are sometimes grouped together into a single embodiment, figure, or description thereof. However, this disclosed approach should not be interpreted as reflecting an intention that the claimed disclosure requires more features than those expressly recited in each claim. Rather, as reflected in the claims below, the disclosed aspects consist of fewer than all the features of the individual embodiments disclosed above. Accordingly, the claims that follow the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present disclosure.

[0101] Those skilled in the art will appreciate that the modules, units, or components of the devices in the examples disclosed herein may be arranged in the device described in the embodiment, or alternatively may be located in one or more devices different from the devices in the examples. The modules in the foregoing examples may be combined into one module or further divided into multiple submodules.

[0102] Those skilled in the art will appreciate that the modules in the devices in the embodiments may be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments may be combined into one module or unit or component, and in addition may be divided into multiple submodules or subunits or subcomponents. All features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed herein may be combined in any combination, except that at least some of such features and / or processes or units are mutually exclusive. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature providing the same, equivalent or similar purpose.

[0103] Furthermore, those skilled in the art will appreciate that although some embodiments described herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of this disclosure and to form different embodiments. For example, in the claims below, any of the claimed embodiments may be used in any combination.

[0104] In addition, some of the embodiments are described herein as methods or combinations of method elements that can be implemented by a processor of a computer system or by other devices that perform the functions described. Thus, a processor having the necessary instructions for implementing the method or method element forms a device for implementing the method or method element. In addition, the elements described herein of the device embodiments are examples of devices for implementing the functions performed by the elements for the purpose of implementing the disclosed subject matter.

[0105] As used herein, unless otherwise specified, the use of ordinal numbers "first," "second," "third," etc. to describe common objects merely indicates that different instances of similar objects are involved and are not intended to imply that the objects so described must have a given order in time, space, ranking, or in any other manner.

[0106] Although the present disclosure has been described with respect to a limited number of embodiments, it will be apparent to those skilled in the art, having the benefit of the foregoing description, that other embodiments are contemplated within the scope of the disclosure thus described. Furthermore, it should be noted that the language used in this specification has been selected primarily for readability and didactic purposes, and not for the purpose of explaining or limiting the subject matter of the disclosure. Accordingly, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the appended claims. With respect to the scope of the disclosure, what has been disclosed is illustrative and not restrictive, and the scope of the disclosure is defined by the appended claims.

Claims

1. A method for identifying user intent, comprising: Generate the user intention set according to the scenario of the user's answer; Dividing the user intention set into multiple intention groups. When the scenario in which the user answers is a scenario for cleaning clues of the user's intention to buy a car, the intention groups include: an intention group indicating a tendency, an intention group asking a question, and an intention group of a neutral category. Acquire voice data containing the user's answer and convert the voice data into text data; Processing the text data using a first intent recognition model to predict a first intent indicating a user's intent; Processing the text data using a second intent recognition model to predict a second intent indicating a user intent, wherein both the first intent and the second intent belong to a user intent set, and the user intent set includes multiple intent groups; Determining, based at least on a relationship between the first intent and each intent group, and the second intent, an intent corresponding to each intent group; The intents corresponding to the intent groups are integrated to obtain the user intent of the user's answer; The first intention predicted by the first intention recognition model is the intention with the largest probability value among all intentions; The second intent predicted by the second intent recognition model is a plurality of intents whose probability values ​​are greater than a threshold among all intents; The determining, based at least on the relationship between the first intent and each intent group, and the second intent, of the intent corresponding to each intent group includes: Setting a corresponding maximum number of retained intents for each intent group, wherein the first predetermined number is the maximum number of retained intents minus 1, the second predetermined number is the maximum number of retained intents, and the maximum number of retained intents for the intent group representing the tendency is 1; If the first intent belongs to the intent group, sequentially selecting a first predetermined number of intents belonging to the intent group from the second intent as the intent corresponding to the intent group, including: if the first intent belongs to the intent group and the maximum number of retained intents is 1, then there is no need to select an intent from the second intent; If the first intent does not belong to the intent group, sequentially selecting a second predetermined number of intents belonging to the intent group from the second intent as the intent corresponding to the intent group; The method of fusing the intents corresponding to the intent groups to obtain the user intent of the user's answer includes: The union of the intents corresponding to the intent groups is taken as the final user intent of the user's answer.

2. The method as claimed in claim 1, wherein the union of the multiple intent groups is the user intent set, and the intent groups are mutually disjoint.

3. The method according to claim 1, wherein The determining, based at least on the relationship between the first intent and each intent group, and the second intent, of the intent corresponding to each intent group further includes: In descending order of probability values ​​of the respective intents in the second intents, a first predetermined number or a second predetermined number of intents belonging to the corresponding intent group are selected from the second intents.

4. The method according to any one of claims 1 or 2, further comprising the step of training the first intent recognition model: Obtain the text data of the user's answer and mark the unique intention of each text data as the first training sample; Inputting the text data into an initial first intent recognition model to predict the probability value of each intent, and taking the intent with the largest probability value as the predicted first intent; The first intent recognition model is trained using the labeled data in the first training sample and the predicted first intent, until the training is completed to obtain a trained first intent recognition model.

5. The method according to any one of claims 1 or 2, further comprising the step of training the second intent recognition model: Obtaining text data of the user's answer and marking an intent set of each text data, wherein the intent set includes multiple intents as a second training sample; Inputting the text data into an initial second intent recognition model to predict a probability value belonging to each intent, and taking the intent with a probability value greater than a threshold as the predicted second intent; The second intent recognition model is trained using the labeled data in the second training sample and the predicted second intent, until the training is completed to obtain a trained second intent recognition model.

6. The method of claim 1, wherein: The intention groups indicating the tendency include at least the following intentions: intention to have already bought a car, intention to buy a car, affirmative intention, and negative intention; The intention grouping of the question asking includes at least the following intentions: intention to ask about price, intention to ask about identity; The neutral category of intentions includes at least the following intentions: intention to hang up, intention to be temporarily inconvenient to communicate, and other intentions.

7. A user intention recognition system comprising: a text data acquisition unit adapted to generate the user intention set according to the scenario of the user's answer, divide the user intention set into a plurality of intention groups, acquire voice data containing the user's answer, and convert the voice data into text data; an intent prediction unit, adapted to process the text data using a first intent recognition model to predict a first intent indicating a user intent, and process the text data using a second intent recognition model to predict a second intent indicating the user intent, wherein the first intent and the second intent both belong to a user intent set, and the user intent set includes multiple intent groups; and the first intent predicted by the first intent recognition model is an intent with the greatest probability value among all intents; The second intent predicted by the second intent recognition model is a plurality of intents whose probability values ​​are greater than a threshold among all intents; an intention determination unit, adapted to determine an intention corresponding to each intention group based on at least a relationship between the first intention and each intention group, and the second intention; It is also suitable for fusing the intentions corresponding to each intention grouping to obtain the user intention answered by the user, and determining the intention corresponding to each intention grouping based at least on the relationship between the first intention and each intention grouping, and the second intention, including: setting the corresponding maximum number of reserved intentions for each intention group, wherein the first predetermined number is the maximum number of reserved intentions minus 1, and the second predetermined number is the maximum number of reserved intentions; if the first intention belongs to the intention group, then selecting a first predetermined number of intentions belonging to the intention group from the second intention in sequence as the intention corresponding to the intention group, including: if the first intention belongs to the intention group and the maximum number of reserved intentions is 1, there is no need to select an intention from the second intention; if the first intention does not belong to the intention group, then selecting a second predetermined number of intentions belonging to the intention group from the second intention in sequence as the intention corresponding to the intention group; fusing the intentions corresponding to each intention group to obtain the user intention answered by the user, including: taking the union of the intentions corresponding to each intention group as the final user intention of the user answer.

8. A computing device comprising: one or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, which, when executed by a computing device, cause the computing device to perform the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • User intention recognition method and device for intelligent voice robot and electronic equipment

    CN112100339A

  • Multi-intention identification method and device, equipment and medium

    CN112541079A