A method and device for determining user intent

By using benchmarks and third-party intent recognition models in the intelligent voice interaction platform, allowing third-party users to customize intent recognition models, solving the problem of lack of scalability and poor user experience in the prior art, and achieving higher intent recognition accuracy and user experience optimization.

CN114694645BActive Publication Date: 2025-06-13HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011628131.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-31
Publication Date
2025-06-13
Estimated Expiration
2040-12-31

AI Technical Summary

Technical Problem

The existing intelligent voice interaction platform lacks scalability when identifying user intentions and cannot effectively support third-party platforms to customize voice intentions, resulting in poor user experience.

Method used

By using at least one benchmark intent recognition model and at least one third-party intent recognition model, the third-party intent recognition model is trained based on the benchmark intent recognition model of the same skill category and its model training data, allowing third-party users to train a custom intent recognition model based on the preset skill category.

Benefits of technology

It enhances the scalability of customized voice intent of third-party platforms, improves the accuracy and user experience of user intention recognition, and reduces the need for users to wake up words in memory skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114694645B_ABST
    Figure CN114694645B_ABST
Patent Text Reader

Abstract

The present application relates to a method and apparatus for determining user intent, and relates to the natural language understanding technology in the field of artificial intelligence. The method includes: obtaining a speech text corresponding to a speech signal; inputting the speech text into at least one benchmark intent recognition model and at least one third-party intent recognition model respectively, outputting a first intent set through the at least one benchmark intent recognition model, and outputting a second intent set through the at least one third-party intent recognition model, wherein the third-party intent recognition model is set to be trained based on the benchmark intent recognition model of the same skill category and its model training data; determining the intent of the speech text according to the confidence of the intent in the first intent set and the confidence of the intent in the second intent set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence voice interaction technology, and particularly to a method and device for determining user intent. Background Art

[0002] In recent years, intelligent speech interaction technology has developed rapidly. Based on technologies such as speech recognition, speech synthesis, and natural language understanding, intelligent speech interaction technology can endow products with an intelligent human-computer interaction experience of "being able to listen, speak, and understand you" for users in various practical application scenarios.

[0003] Currently, intelligent speech interaction platforms often need to cooperate with multiple third-party platforms to provide users with rich voice skills. Typically, the cooperating third-party platforms mainly include merchants, music radio platforms, weather information platforms, and so on. Since the number of third-party platforms is large and many third-party platforms belong to the same type, it has become very important to accurately identify which skill of which platform the user wants to trigger. Usually, intelligent speech interaction platforms only support opening skills with skill wake-up words to third-party platforms, and these skills can only be recalled through the user's voice text with a clear skill wake-up word. In one example, the skill wake-up word for playing music can be set as "play music". Then, if the user wants to listen to a certain song, they need to first say the skill wake-up word "play music", and then say the name of the song. Since there are many voice skills involved in intelligent speech interaction platforms, the method of waking up skills using skill wake-up words has relatively high requirements for users, and users cannot remember too many skill wake-up words. Furthermore, triggering voice skills without a skill wake-up word has become a more popular voice interaction method for users. Triggering voice skills without a skill wake-up word means triggering voice skills without having to say the skill wake-up word. For example, in the above example, the user does not need to first say the skill wake-up word "play music", and the user can directly say "play XY" (XY is the name of the song) to trigger the skill of playing music. In related technologies, intelligent speech interaction platforms can often develop multiple preset intents, and these preset intents are often not modifiable. If a third-party platform supports one or some of these preset intents, the corresponding preset intent can be referenced. In this way, when the user's voice hits one of the preset intents and the preset intent corresponds to the voice skills of multiple third-party platforms, the user can be confirmed which third-party platform's voice skill to use. In the method of related technologies, third-party platforms can only reference the preset intents already defined by intelligent speech interaction platforms, and cannot expand the existing preset intents, resulting in poor scalability.

[0004] Therefore, there is an urgent need in related technologies for a way to provide better extensible custom voice intents for third-party platforms. Summary of the Invention

[0005] In view of this, a method and device for determining user intent are proposed.

[0006] In a first aspect, an embodiment of the present application provides a method for determining user intent.

[0007] According to the first aspect, in a first possible implementation manner, it includes:

[0008] Obtain the speech text corresponding to the speech signal;

[0009] Input the speech text into at least one benchmark intent recognition model and at least one third-party intent recognition model respectively. The at least one benchmark intent recognition model outputs a first intent set, and the at least one third-party intent recognition model outputs a second intent set. Wherein, the third-party intent recognition model is trained based on the benchmark intent recognition model of the same skill category and its model training data;

[0010] Determine the intent of the speech text according to the confidence of the intent in the first intent set and the confidence of the intent in the second intent set.

[0011] The method for determining user intent provided by each embodiment of the present application can use at least one benchmark intent recognition model and at least one third-party intent recognition model to recognize the user's speech text. Among them, the third-party intent recognition model is set to be trained based on the benchmark intent recognition model of the same skill category and its model training data. Thus, in the embodiments of the present application, conditions for training an intent recognition model can be provided to third-party users. Third-party users can train their own intent recognition models based on the benchmark intent recognition models of preset skill categories, enhancing the expandability of third-party users' custom intents. On the other hand, from the user's perspective, using multiple intent recognition models to recognize the user's speech text can help the user recall intents with higher accuracy and optimize the user experience.

[0012] According to the first possible implementation manner of the first aspect, the third-party intent recognition model is trained in the following manner:

[0013] Obtain the benchmark intent recognition model of the preset skill category and the model training data of the benchmark intent recognition model. The model training data at least includes a plurality of preset intents and the benchmark sample data and benchmark model parameters respectively corresponding to the plurality of preset intents;

[0014] Obtain third-party sample data matching the preset skill category;

[0015] Train the baseline intent recognition model using the model training data and the third-party sample data to generate the third-party intent recognition model.

[0016] This embodiment provides a specific method for training a third-party intent recognition model. The third-party intent recognition model is developed based on the original baseline intent recognition model, which can not only prevent the third-party intent recognition model from polluting the intent of the baseline intent recognition model, but also reduce the difficulty of developing third-party intents and improve the parsing ability of the third-party intent recognition model based on the training and development of the mature baseline intent recognition model.

[0017] According to the second possible implementation manner of the first aspect, the obtaining of the third-party sample data matching the preset skill category includes:

[0018] Obtain the third-party intents added by the third-party user and the third-party sample data corresponding to the third-party intents, where the third-party intents match the preset skill category, or,

[0019] Obtain the sample data added by the third-party user on the basis of the baseline sample data corresponding to the preset intent.

[0020] This embodiment provides a way for the third-party user to provide sample data. On the one hand, the third-party user can add third-party intents. On the other hand, the third-party user can add sample data for existing intents. Thus, it can be seen that the third-party user can not only expand intents, but also enrich the sample data of existing intents, making the trained third-party intent recognition model more in line with the business characteristics of the third-party user.

[0021] According to the third possible implementation manner of the first aspect, the training of the baseline intent recognition model using the model training data and the third-party sample data to generate the third-party intent recognition model includes:

[0022] Obtain the user identification of the third-party user;

[0023] Associate the user identification with the third-party sample data;

[0024] Train the baseline intent recognition model using the model training data and the third-party sample data associated with the user identification to generate the third-party intent recognition model.

[0025] In the embodiments of the present application, the user identification of the third-party user can be uniformly substituted into the sample data for training. On the one hand, substituting the user identification during the training process adapts to the habit that the user prefers to use the merchant's user identification in the voice command. On the other hand, it avoids adding the user identification to each third-party sample data provided by the third-party user, increasing information redundancy.

[0026] According to the fourth possible implementation manner of the first aspect, the user identifier includes at least one of the brand name, APP name, and product name corresponding to the third-party user.

[0027] The embodiments of the present application provide various possible user identifiers.

[0028] According to the fifth possible implementation manner of the first aspect, determining the intent of the speech text according to the confidence of the intents in the first intent set and the confidence of the intents in the second intent set includes:

[0029] When it is determined that the confidence of the intents included in the first intent set is less than or equal to a first preset threshold, and the confidence of the intents included in the second intent set is greater than a second preset threshold, taking the intent with the highest confidence in the second intent set as the intent of the speech text; or,

[0030] When it is determined that the confidence of the intents included in the second intent set is less than or equal to the second preset threshold, taking the intent with the highest confidence in the first intent set as the intent of the speech text; or,

[0031] When it is determined that the confidence of the intents included in the first intent set is greater than or equal to the first preset threshold, and the confidence of the intents included in the second intent set is greater than the second preset threshold, taking the intent with the highest confidence in the first intent set and the second intent set as the intent of the speech text.

[0032] The embodiments of the present application provide various ways to determine the user intent, improving the recall rate of intent recognition.

[0033] According to the sixth possible implementation manner of the first aspect, the first preset threshold and the second preset threshold are set to match the corresponding skill categories respectively.

[0034] In this embodiment, different confidence thresholds can be set for different skill categories respectively to adapt to the confidence characteristics corresponding to different skill categories.

[0035] In a second aspect, an embodiment of the present application provides a method for generating an intent recognition model.

[0036] According to the second aspect, in the first possible implementation manner, it includes:

[0037] Obtain the preset skill category selected by the third-party user;

[0038] Obtain the benchmark intent recognition model corresponding to the preset skill category and its model training data;

[0039] Obtain third - party sample data from the third - party user that matches the preset skill category;

[0040] Use the model training data and the third - party sample data to train the benchmark intent recognition model to generate a third - party intent recognition model, where the third - party intent recognition model is an intent recognition model corresponding to the third - party user.

[0041] According to the first possible implementation manner of the second aspect, the obtaining of the third - party sample data that matches the preset skill category includes:

[0042] Obtain the third - party intent added by the third - party user and the third - party sample data corresponding to the third - party intent, where the third - party intent matches the preset skill category, or,

[0043] Obtain the sample data added by the third - party user based on the benchmark sample data corresponding to the preset intent.

[0044] According to the second possible implementation manner of the second aspect, the using of the model training data and the third - party sample data to train the benchmark intent recognition model to generate the third - party intent recognition model includes:

[0045] Obtain the user identifier of the third - party user;

[0046] Associate the user identifier with the third - party sample data;

[0047] Use the model training data and the third - party sample data associated with the user identifier to train the benchmark intent recognition model to generate the third - party intent recognition model.

[0048] According to the third possible implementation manner of the second aspect, the user identifier includes at least one of the brand name, APP name, and product name corresponding to the third - party user.

[0049] According to the fourth possible implementation manner of the second aspect, after generating the third - party intent recognition model, it further includes:

[0050] Obtain the speech text corresponding to the speech signal;

[0051] Input the speech text into at least one benchmark intent recognition model and at least one third - party intent recognition model respectively. The at least one benchmark intent recognition model outputs a first intent set, and the at least one third - party intent recognition model outputs a second intent set;

[0052] Determine the intent of the speech text according to the confidence of the intents in the first intent set and the confidence of the intents in the second intent set.

[0053] In a third aspect, embodiments of the present application provide a device or system for determining a user's intent.

[0054] According to the third aspect, in a first possible implementation manner, the device or system (the system may be a software platform, such as the open platform mentioned below) may include:

[0055] A speech recognition module, configured to obtain a speech text corresponding to a speech signal;

[0056] A dialogue management module, configured to input the speech text into at least one benchmark intent recognition model and at least one third-party intent recognition model respectively. The at least one benchmark intent recognition model outputs a first intent set, and the at least one third-party intent recognition model outputs a second intent set. Among them, the third-party intent recognition model is trained based on the benchmark intent recognition model of the same skill category and its model training data; and is configured to determine the intent of the speech text according to the confidence of the intents in the first intent set and the confidence of the intents in the second intent set.

[0057] According to the first possible implementation manner of the third aspect, the third-party intent recognition model is trained in the following manner:

[0058] Obtain a benchmark intent recognition model of a preset skill category and the model training data of the benchmark intent recognition model. The model training data at least includes a plurality of preset intents and the benchmark sample data and benchmark model parameters respectively corresponding to the plurality of preset intents;

[0059] Obtain third-party sample data matching the preset skill category;

[0060] Train the benchmark intent recognition model using the model training data and the third-party sample data to generate the third-party intent recognition model.

[0061] According to the second possible implementation manner of the third aspect, the obtaining of the third-party sample data matching the preset skill category includes:

[0062] Obtain third-party intents added by a third-party user and the third-party sample data corresponding to the third-party intents. The third-party intents match the preset skill category, or,

[0063] Obtain sample data added by a third-party user based on the benchmark sample data corresponding to the preset intent.

[0064] According to the third possible implementation manner of the third aspect, the training of the benchmark intent recognition model using the model training data and the third-party sample data to generate the third-party intent recognition model includes:

[0065] Obtain the user identifier of the third-party user;

[0066] Associate the user identifier with the third-party sample data;

[0067] Use the model training data and the third-party sample data associated with the user identifier to train the benchmark intent recognition model to generate the third-party intent recognition model.

[0068] According to the fourth possible implementation manner of the third aspect, the user identifier includes at least one of the brand name, APP name, and product name corresponding to the third-party user.

[0069] According to the fifth possible implementation manner of the third aspect, the determining of the intent of the speech text according to the confidence levels of the intents in the first intent set and the confidence levels of the intents in the second intent set includes:

[0070] When it is determined that the confidence levels of all the intents included in the first intent set are less than or equal to a first preset threshold and the confidence levels of all the intents included in the second intent set are greater than a second preset threshold, use the intent with the highest confidence level in the second intent set as the intent of the speech text; or,

[0071] When it is determined that the confidence levels of all the intents included in the second intent set are less than or equal to the second preset threshold, use the intent with the highest confidence level in the first intent set as the intent of the speech text; or,

[0072] When it is determined that the confidence levels of all the intents included in the first intent set are greater than or equal to the first preset threshold and the confidence levels of all the intents included in the second intent set are greater than the second preset threshold, use the intent with the highest confidence level in the first intent set and the second intent set as the intent of the speech text.

[0073] According to the sixth possible implementation manner of the third aspect, the first preset threshold and the second preset threshold are set to match the corresponding skill categories respectively.

[0074] Fourth aspect, an embodiment of the present application provides a device or system for generating an intent recognition model (the system can be a software system).

[0075] According to the fourth aspect, in the first possible implementation manner, it includes:

[0076] A skill category acquisition module for acquiring a preset skill category selected by a third-party user;

[0077] A model acquisition module for acquiring a benchmark intent recognition model corresponding to the preset skill category and its model training data;

[0078] A sample acquisition module for acquiring third-party sample data from the third-party user that matches the preset skill category;

[0079] A model generation module for training the benchmark intent recognition model using the model training data and the third-party sample data to generate a third-party intent recognition model, where the third-party intent recognition model is an intent recognition model corresponding to the third-party user.

[0080] According to the first possible implementation manner of the fourth aspect, acquiring the third-party sample data that matches the preset skill category includes:

[0081] Acquiring third-party intents added by the third-party user and third-party sample data corresponding to the third-party intents, where the third-party intents match the preset skill category, or,

[0082] Acquiring sample data added by the third-party user based on the benchmark sample data corresponding to the preset intent.

[0083] According to the second possible implementation manner of the fourth aspect, training the benchmark intent recognition model using the model training data and the third-party sample data to generate the third-party intent recognition model includes:

[0084] Acquiring the user identifier of the third-party user;

[0085] Associating the user identifier with the third-party sample data;

[0086] Training the benchmark intent recognition model using the model training data and the third-party sample data associated with the user identifier to generate the third-party intent recognition model.

[0087] According to the third possible implementation manner of the fourth aspect, the user identifier includes at least one of the brand name, APP name, and product name corresponding to the third-party user.

[0088] According to the fourth possible implementation manner of the fourth aspect, it further includes:

[0089] A speech recognition module for acquiring the speech text corresponding to the speech signal;

[0090] A dialogue management module, configured to input the speech text into at least one benchmark intent recognition model and at least one third-party intent recognition model respectively, output a first intent set through the at least one benchmark intent recognition model, and output a second intent set through the at least one third-party intent recognition model; and, configured to determine the intent of the speech text according to the confidence levels of the intents in the first intent set and the confidence levels of the intents in the second intent set.

[0091] In a fifth aspect, an embodiment of the present application provides a terminal device, including:

[0092] A processor; a memory for storing processor-executable instructions;

[0093] Wherein, the processor is configured to execute the instructions such that the terminal device implements the method according to one or several of the above first / second aspects or various possible implementations of the first / second aspects.

[0094] In a sixth aspect, an embodiment of the present application provides a non-volatile computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions are executed by a processor to implement the method according to one or several of the above first / second aspects or various possible implementations of the first / second aspects.

[0095] In a seventh aspect, an embodiment of the present application provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, the processor in the electronic device executes the method according to one or several of the above first / second aspects or various possible implementations of the first / second aspects.

[0096] In an eighth aspect, an embodiment of the present application provides a chip, which includes at least one processor, and the processor is used to run computer programs or computer instructions stored in a memory to execute the method according to any possible implementation of the above aspects.

[0097] Optionally, the chip may further include a memory for storing computer programs or computer instructions. Optionally, the chip may further include a communication interface for communicating with other modules outside the chip.

[0098] Optionally, one or more chips may form a chip system. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] The drawings included in and constituting a part of the specification illustrate exemplary embodiments, features, and aspects of the present application together with the specification, and are used to explain the principles of the present application.

[0100] Figure 1 Shows a scene example diagram according to an embodiment of the present application.

[0101] Figure 2 Shows a scene example diagram according to an embodiment of the present application.

[0102] Figure 3 Shows a scene example diagram according to an embodiment of the present application.

[0103] Figure 4 Shows a scene example diagram according to an embodiment of the present application.

[0104] Figure 5 Shows a flowchart of a method for determining user intention according to an embodiment of the present application.

[0105] Figure 6 Shows a flowchart of a method for determining user intention according to an embodiment of the present application.

[0106] Figure 7 Shows a schematic structural diagram of a terminal device according to an embodiment of the present application. Detailed implementation manners

[0107] The following will describe various exemplary embodiments, features, and aspects of the present application in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0108] The special term "exemplary" here means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" here is not necessarily to be construed as superior to or better than other embodiments.

[0109] In addition, for a better description of the present application, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present application can also be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present application.

[0110] To clearly show the method for determining user intention provided by the embodiments of the present application, the technical solution will be described below through a specific application scenario.

[0111] In this exemplary scenario, the user's search intention is to search for delicious dumplings nearby using XX, where XX can be the name of a certain application. Thus, in step 1, the user sends a voice signal through the voice assistant. The voice assistant can, for example, include devices with a microphone and a communication module, such as smartphones, speakers, computers, smart wearable devices, and so on. After receiving the voice signal, the voice assistant can send the voice signal to the speech recognition module. The speech recognition module is used to recognize the speech text corresponding to the voice signal. For example, the speech recognition module recognizes that the speech text corresponding to the voice signal is "Search for the nearest delicious dumplings using XX". In step 2, the speech recognition module can send the speech text to the dialogue management module. The dialogue management module is used to determine the user intention corresponding to the voice signal. Based on the technical solution provided by the embodiments of the present application, in steps 3 and 4, the dialogue management module can send the speech text to at least one baseline intention recognition model and at least one third-party intention recognition model. The intention recognition model and the third-party intention recognition model can, for example, include NLU (Natural Language Understanding) models.

[0112] In the embodiments of the present application, the intelligent voice interaction platform can provide baseline intention recognition models of multiple preset skill categories, such as Figure 1 and Figure 2 As shown, the baseline intention recognition models can include a food category baseline intention recognition model, a music category baseline intention recognition model, a hotel category baseline intention recognition model, a train ticket category baseline intention recognition model, a weather category baseline intention recognition model, a flight ticket category baseline intention recognition model, and so on. The preset skill categories can include at least one of the following: food category (or referred to as querying food category, similar understandings can be made for other similar parts in this article), music category, hotel category, train ticket category, weather category, flight ticket category, etc. Based on the baseline intention recognition models of multiple preset skill categories provided by the intelligent voice interaction platform, a third-party platform or a third-party user (i.e., the developer or maintainer of the third-party platform) can develop a custom third-party intention recognition model. For example, for third-party users, such as Party A and Party B providing food services, they can train their own food category intention recognition models based on the food category baseline intention recognition model without developing a skill triggering method with a custom skill wake-up word (such as not developing a skill wake-up word corresponding to the food category). For Party C providing travel services, the service items involved include food, hotels, train tickets, flight tickets, etc. Based on this, Party C can train its own intention recognition models for multiple service items on the basis of multiple baseline intention recognition models such as food, hotels, train tickets, and flight tickets.

[0113] Figure 3 A schematic diagram showing the training of a third-party intent recognition model based on a benchmark intent recognition model is shown. As Figure 3 shown, a third-party platform can provide sample data corresponding to a preset skill category. For example, Party A or Party B can provide sample data for the food category, and Party C can provide sample data for the food category, hotel category, train ticket category, and flight ticket category. In this way, after obtaining the sample data provided by the third party, on the basis of the benchmark intent recognition model, the sample data provided by the third party can be added to train the benchmark intent recognition model to generate the third-party intent recognition model. In the embodiments of the present application, during the training process, the model training data of the benchmark intent recognition model can be retained. The model training data can include: a training sample data set, model parameters, a validation sample data set, and so on. In this way, after adding the sample data provided by the third party, the benchmark intent recognition model can be fine-tuned to generate the third-party intent recognition model. In the embodiments of the present application, based on the manner of generating the third-party intent recognition model, the same preset skill category can correspond to multiple third-party intent recognition models. As Figure 1 shown, the third-party intent recognition models for the food category can include the food category intent recognition model of Party A, the food category intent recognition model of Party B, the food category intent recognition model of Party C, and so on. The third-party intent recognition models for the weather category can include the weather category intent recognition model of Party C and the weather category intent recognition model of Party F.

[0114] In steps 5 and 6, after inputting the speech text "Search for the nearest delicious dumplings with XX" into at least one benchmark intent recognition model, a first intent set is output by the at least one benchmark intent recognition model. The first intent set can include multiple intents and the confidence levels respectively corresponding to each intent. After inputting the speech text into at least one third-party intent recognition model, a second intent set is output by the at least one third-party intent recognition model. The second intent set can also include multiple intents and the confidence levels respectively corresponding to each intent. The intents in the first intent set can include a predetermined number of intents with the highest confidence levels output by the at least one benchmark intent recognition model. The intents in the second intent set can include a predetermined number of intents with the highest confidence levels output by the at least one third-party intent recognition model.

[0115] After obtaining the first intent set and the second intent set, the dialogue management module may determine the final user intent from the first intent set and the second intent set. In an embodiment of the present application, when it is determined that the confidence levels of the intents included in the first intent set are all less than or equal to a first preset threshold, and the confidence levels of the intents included in the second intent set are all greater than a second preset threshold, the intent with the highest confidence level in the second intent set is taken as the intent of the speech text. In another embodiment of the present application, when it is determined that the confidence levels of the intents included in the second intent set are all less than or equal to the second preset threshold, the intent with the highest confidence level in the first intent set is taken as the intent of the speech text. In another embodiment of the present application, when it is determined that the confidence levels of the intents included in the first intent set are all greater than or equal to the first preset threshold, and the confidence levels of the intents included in the second intent set are all greater than the second preset threshold, the intent with the highest confidence level in the first intent set and the second intent set is taken as the intent of the speech text. For example, the dialogue management module determines that the final user intent is to recommend food using the XX APP according to the above method.

[0116] As Figure 1 shown, after the dialogue management module determines the user intent corresponding to the speech text, in step 7, the user intent can be input into the intent implementation configuration module. The intent implementation configuration module can determine the implementation method of the user intent. For example, it can provide an API or a jump link connecting to a third-party platform to the dialogue management module. After receiving the API or the jump link, in step 9, the dialogue management module can display a relevant page in the APP corresponding to the third-party platform.

[0117] Figure 4 The functional modules involved in the embodiments of the present application may include:

[0118] 1) In terms of hardware, it may include a microphone or a microphone array.

[0119] The microphone or the microphone array may be disposed in the voice assistant device. The microphone or the microphone array can not only acquire sound signals, but also enhance the sound in the direction of the sound source and suppress the noise in the non-sound source direction.

[0120] In addition, by means of the cooperation of a camera + microphone array, directional noise cancellation of sound can be achieved.

[0121] 2) Local processing may include a signal processing module.

[0122] The signal processing module not only processes the acquired voice signals, such as amplifying and filtering them, but also can determine the angle of the sound source after determining the position of the sound source, and then control the sound pickup of the microphone or microphone array to achieve directional noise cancellation.

[0123] 3) Cloud processing, that is, implemented in the cloud. Of course, it can also be local processing, which can be determined according to the processing power of the device itself and the usage environment, etc. Of course, if it is implemented in the cloud, by leveraging big data to update and adjust the algorithm model, the accuracy of speech recognition, natural speech understanding, and dialogue management can be effectively improved.

[0124] Cloud processing can involve implementing the functions of at least one of the following modules in the cloud: speech recognition module, dialogue management module, intent recognition module, and intent implementation configuration module, etc. Among them,

[0125] The speech recognition module is mainly used to recognize the speech text of the acquired voice signals. For example, if a piece of speech is acquired and its meaning needs to be understood, then it is necessary to first know the specific text content of this piece of speech. This process requires the speech recognition module to convert the voice signal into speech text.

[0126] For a machine, words are still just the words themselves. To determine the meaning expressed by the words, it is necessary to determine the natural meaning corresponding to the speech text, so as to recognize the intention of the user's speech. The purpose of the dialogue management module is to achieve effective communication with the user to obtain the information required for performing operations. In the embodiments of the present application, the dialogue management module can be used to call the intent recognition module to obtain the intent set and the confidence of the intent, and make a decision on the final intent according to the intent set and the confidence of the intent therein, and call the intent implementation configuration module to obtain the implementation method of the intent according to the intent and the skill category to which the intent belongs.

[0127] For the functions corresponding to one or more of the specific speech recognition module, dialogue management module, intent recognition module, and intent implementation configuration module, they can be processed in the cloud (i.e., implemented through a server, such as a cloud server), or can be processed locally (e.g., implemented on the terminal device, rather than through a server), which can be determined according to the processing power of the device itself and the usage environment, etc. Of course, if it is processed in the cloud, by leveraging big data to update and adjust the algorithm model, the accuracy of speech recognition, natural speech understanding, and dialogue management can be effectively improved.

[0128] The method for determining the user's intention according to the present application will be described in detail below with reference to the accompanying drawings. Figure 5It is a schematic flowchart of a method according to an embodiment of the method for determining user intention provided by this application. Although this application provides method operation steps as shown in the following embodiments or drawings, more or fewer operation steps may be included in the method based on routine or non-creative labor. In steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided by the embodiments of this application. When the method actually determines the user intention process or the device executes, it can be executed in the method order shown in the embodiments or drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing).

[0129] As Figure 5 shown, in an embodiment of this application, the method may include:

[0130] S501: Obtain the speech text corresponding to the speech signal.

[0131] In the embodiments of this application, the user can send a speech signal through a voice assistant. The voice assistant may include, for example, a device with a microphone and a communication module, such as a smart phone, a speaker, a computer, a smart wearable device, etc. After receiving the speech signal, the voice assistant can send the speech signal to a local or cloud-based speech recognition module. The speech recognition module can recognize the speech text corresponding to the speech signal. The speech recognition module may include, for example, an Automatic Speech Recognition (ASR) module.

[0132] S502: Input the speech text into at least one benchmark intention recognition model and at least one third-party intention recognition model respectively. The first intention set is output by the at least one benchmark intention recognition model, and the second intention set is output by the at least one third-party intention recognition model, where the third-party intention recognition model is trained based on the benchmark intention recognition model of the same skill category and its model training data.

[0133] In the embodiments of this application, the third-party intention recognition model is set to be trained based on the benchmark intention recognition model of the same skill category and its model training data. The skill category may include, for example, music skills, video skills, food skills, weather skills, flight ticket skills, etc. That is to say, if it is necessary to train a third-party intention recognition model for food, it needs to be trained based on the benchmark intention recognition model for food. If it is necessary to train a third-party intention recognition model for music, it needs to be trained based on the benchmark intention recognition model for music. In the embodiments of this application, benchmark intention recognition models of multiple skill categories can be provided to third-party users.

[0134] In one embodiment of the present application, the third-party intent recognition model can be trained in the following manner:

[0135] SS1: Obtain a benchmark intent recognition model for a preset skill category and the model training data of the benchmark intent recognition model. The model training data at least includes a plurality of preset intents and the benchmark sample data and benchmark model parameters respectively corresponding to the plurality of preset intents.

[0136] SS3: Obtain third-party sample data that matches the preset skill category.

[0137] SS5: Use the model training data and the third-party sample data to train the benchmark intent recognition model to generate the third-party intent recognition model.

[0138] The following combines Figure 6 to illustrate the above embodiment. As Figure 6 shown, an open platform for training an intent recognition model can be provided to third-party users (which can also be understood as a third-party platform here). The open platform can include the intelligent voice interaction platform. Third-party users can train a third-party intent recognition model on the open platform. Based on this, as Figure 6 shown in step 1, the third party first needs to determine the skill category of the intent recognition model to be trained. In step 2, after the open platform obtains the preset skill category determined by the third party, it can obtain the benchmark intent recognition model for the preset skill category and the model training data of the benchmark intent recognition model. In the embodiment of the present application, an area for storing the benchmark intent recognition model and the third-party intent recognition model can be set. Based on this, the open platform can obtain the required benchmark intent recognition model and its model training data from this area. The model training data can at least include a plurality of preset intents and the benchmark sample data and benchmark model parameters respectively corresponding to the plurality of preset intents. Table 1 shows several preset intents included in the benchmark intent recognition model for the food category and related information. As shown in Table 1, the information of the preset intent can include an intent unique identifier, an intent name, an implementation method, an operation, etc. The benchmark sample data corresponding to the preset intent can include the corpus used to train the preset intent, etc. For example, for "order takeout", the corresponding corpus can include "help me order takeout", "I want to order fried chicken", "help me order a cup of coffee", etc. Of course, the benchmark sample data can also include at least one piece of information such as the slot identifier, slot dictionary, slot value corresponding to the benchmark sample data. For example, for the corpus "What's the weather like in Shanghai today?", the corresponding slot and slot value can be marked as "city = Shanghai", "time = today". By annotating the corresponding slot information in the sample data, the slot prediction capabilities of the benchmark intent recognition model and the third-party intent recognition model can be enhanced.

[0139] Table 1 Preset Intent Information Table of the Benchmark Intent Recognition Model for Food Category

[0140] Intention Identifier Intention Name Implementation Method Operation ORDER_TAKEOUT Order takeout Deeplink Modify SEARCH_CATE Search for food RESTful API Modify BOOK_RESTAURANT Find a restaurant RESTful API Modify

[0141] In the embodiments of the present application, as Figure 6 shown in step 3, a third-party user can provide third-party sample data to the open platform. In an embodiment of the present application, based on the multiple preset intents included in the benchmark intent recognition model, the third-party user can add custom third-party intents and third-party sample data corresponding to the third-party intents, where the third-party intents need to match the preset skill categories. In one example, based on the preset intents provided by the benchmark intent recognition model for the food category shown in Table 1, a third-party user can add other custom preset intents for the food category. For example, add a third-party intent of "search for food discount coupons". As shown in Table 2 below, the formed benchmark intent recognition model for the food category of third parties can include the following intents, where the added third-party intents can be set with user-defined intent identifiers, intent names, and implementation methods. Of course, after the third-party user adds third-party intents, third-party sample data corresponding to the third-party intents also needs to be provided. For example, for the third-party intent of "search for food discount coupons", the added third-party sample data can include, for example, "Are there any coffee coupons?" and "Is there a discount on fried chicken of XX brand?" etc.

[0142] Table 2 Intent Information Table of the Benchmark Intent Recognition Model for Food Category of Third Parties

[0143]

[0144]

[0145] In the embodiments of the present application, it can be set that the third-party user cannot delete the preset intents included in the benchmark intent recognition model. In this way, not only can it prevent the intents of the benchmark intent recognition model of the platform from being polluted by the third-party intent recognition model, but also based on the training and development of the mature benchmark intent recognition model, it can reduce the difficulty of third-party developed intents and improve the parsing ability of the third-party intent recognition model. However, in an embodiment of the present application, the third-party user can perform a modification operation on the preset intents. The scope of modification can include, for example, modifying the intent name or adding sample data of the preset intents. In one example, the third-party user finds that the benchmark sample data corresponding to the preset intent of "order takeout" provided by the open platform is not rich enough or does not cover the third-party user's own special products. Based on this, the third-party user can provide some third-party sample data to enrich the quantity of sample data. For example, sample data containing the names of the third-party user's special products can be added.

[0146] It should be noted that the open platform can also review the third-party intents and third-party sample data provided by the third-party users, so that the intents or sample data provided by the third-party users match the selected skill categories. In one example, during the process of training a third-party intent recognition model for the food category, if a third-party user adds a third-party intent of "turn on music", and the open platform determines that this third-party intent is significantly incompatible with the corresponding food skill category during the review, it can provide feedback to the third-party user. Similarly, the open platform can also review the sample data provided by the third-party users, and if it determines that the corresponding sample data is incompatible with the corresponding skill category, it can also provide feedback to the third-party user.

[0147] As Figure 6 shown in step 4 of [], the open platform can use the model training data of the baseline intent recognition model and the third-party sample data to train the baseline intent recognition model to generate the third-party intent recognition model. In an embodiment of the present application, a fine-tuning algorithm can be used to train the baseline intent recognition model. Specifically, only some network layers of the baseline intent recognition model can be adjusted, such as only the last network layer. This method is more suitable for the case where the number of sample data provided by the third-party user is limited. In this way, the third-party intent recognition model can be quickly trained. Of course, in other embodiments, the baseline intent recognition model can also be retrained, especially suitable for the case where the number of sample data provided by the third-party user is large. The present application does not limit this here.

[0148] In an actual application environment, voice commands sent by users (ordinary users, different from the third-party users mentioned in this article) often contain user identifiers of certain third parties (which can also be referred to as third-party platforms or third-party users) here. For example, "order a cup of XX coffee" (where XX is the coffee brand), "search for the nearest delicious dumplings with XX" (where XX is a food application), "open XX music" (where XX is a music playback application), and so on. Based on this, in an embodiment of the present application, during the process of training the third-party intent recognition model, the third-party sample data can be associated with the user identifier of the third-party user. The user identifier includes, for example, at least one of the brand name, APP name, and product name of the third-party user. In an actual application environment, it is difficult for an open platform to require that each sample data provided by a third-party user contains the corresponding user identifier, especially when the number of sample data is large. Based on this, in an embodiment of the present application, the user identifier provided by the third-party user can be obtained first. During the process of training the model, the user identifier can be automatically associated with each piece of the third-party sample data respectively. The association method can include, for example, adding the user identifier to the third-party sample data. Then, the benchmark intent recognition model can be trained using the model training data and the third-party sample data associated with the user identifier to generate the third-party intent recognition model.

[0149] As Figure 6 shown in step 5, after training and generating the third-party intent recognition model, the third-party intent recognition model can be stored in the intent recognition model set. In this way, in the intent recognition model set, for the same skill category, there can be third-party intent recognition models trained and generated by multiple different third-party users.

[0150] Based on the above embodiments of generating the third-party intent recognition model through various trainings, when the speech text is input into at least one benchmark intent recognition model and at least one third-party intent recognition model respectively, a first intent set can be output by the at least one benchmark intent recognition model, and a second intent set can be output by the at least one third-party intent recognition model. The first intent set and the second intent set respectively include multiple recognized intents and the confidence levels of each intent. In the embodiments of the present application, since the number of intents involved is relatively large, the number of intents included in the first intent set and the second intent set can be less than or equal to a preset number. For example, the first intent set may include the M intents with the highest confidence levels recognized by the at least one benchmark intent recognition model. M can be set to 5, 10, 20, etc., and the present application does not limit this here. The second intent set may include the N intents with the highest confidence levels recognized by the at least one third-party intent recognition model. N can be set to 20, 50, 80, etc., and the present application does not limit this here.

[0151] S503: Determine the intent of the speech text according to the confidence levels of the intents in the first intent set and the confidence levels of the intents in the second intent set.

[0152] In the embodiments of the present application, after obtaining the confidence levels of the intents in the first intent set and the confidence levels of the intents in the second intent set, the intent corresponding to the speech text can be determined.

[0153] In an embodiment of the present application, when it is determined that the confidence levels of all the intents included in the first intent set are less than or equal to a first preset threshold, and the confidence levels of all the intents included in the second intent set are greater than a second preset threshold, the intent with the highest confidence level in the second intent set is used as the intent of the speech text. In another embodiment of the present application, when it is determined that the confidence levels of all the intents included in the second intent set are less than or equal to the second preset threshold, the intent with the highest confidence level in the first intent set is used as the intent of the speech text. In another embodiment of the present application, when it is determined that the confidence levels of all the intents included in the first intent set are greater than or equal to the first preset threshold, and the confidence levels of all the intents included in the second intent set are greater than the second preset threshold, the intents with the highest confidence levels in the first intent set and the second intent set are used as the intent of the speech text.

[0154] In the embodiments of the present application, the first preset threshold and the second preset threshold can be used to screen out intents with relatively high confidence and exclude some intents with very low confidence. The first preset threshold and the second preset threshold are set to match the corresponding skill categories respectively. In one example, for intents related to food, the first preset threshold can be set to a, and the second preset threshold can be set to b. For example, a is 0.75 and b is 0.7. On the other hand, for intents related to music, the first preset threshold can be set to c, and the second preset threshold can be set to d. The values of the first preset threshold and the second preset threshold can be determined according to factors such as the performance of the intent recognition model, and the present application does not limit this here.

[0155] In the embodiments of the present application, after determining the user intent corresponding to the voice text, the user intent can be realized. As shown in Table 1 and Table 2, in the embodiments of the present application, the implementation methods of each intent can be predefined, and the implementation methods can include at least one of the following: Deeplink, RESTful API, returning the required text, etc. For example, if it is recognized that the user's intent is to order takeout on XX APP, then the takeout ordering page in XX APP can be displayed in the user interface through the Deeplink method, or a jump link can be displayed in the user interface using the RESTful API method, etc. Of course, the information required by the user can also be directly displayed in the user interface. For example, if the user wants to know the weather information, after obtaining the weather information, the corresponding weather information can be displayed in the user interface in the form of text or a card.

[0156] The method for determining the user intent provided by each embodiment of the present application can use at least one benchmark intent recognition model and at least one third-party intent recognition model to recognize the user's voice text, wherein the third-party intent recognition model is trained based on the benchmark intent recognition model of the same skill category and its model training data. Thus, in the embodiments of the present application, conditions for training the intent recognition model can be provided to third-party users, and third-party users can train their own intent recognition models based on the benchmark intent recognition models of preset skill categories, enhancing the expandability of third-party user-defined intents. On the other hand, from the user's perspective, using multiple intent recognition models to recognize the user's voice text can help the user recall intents with relatively high accuracy and optimize the user experience.

[0157] On the other hand, the present application also provides a method for generating an intent recognition model, including:

[0158] Obtaining a preset skill category selected by a third-party user;

[0159] Obtain the baseline intent recognition model corresponding to the preset skill category and its model training data;

[0160] Obtain third-party sample data from the third-party user that matches the preset skill category;

[0161] Use the model training data and the third-party sample data to train the baseline intent recognition model to generate a third-party intent recognition model, where the third-party intent recognition model is an intent recognition model corresponding to the third-party user.

[0162] Optionally, in an embodiment of the present application, the obtaining of the third-party sample data that matches the preset skill category includes:

[0163] Obtain the third-party intents added by the third-party user and the third-party sample data corresponding to the third-party intents, where the third-party intents match the preset skill category, or,

[0164] Obtain the sample data added by the third-party user based on the baseline sample data corresponding to the preset intent.

[0165] Optionally, in an embodiment of the present application, the using of the model training data and the third-party sample data to train the baseline intent recognition model to generate the third-party intent recognition model includes:

[0166] Obtain the user identifier of the third-party user;

[0167] Associate the user identifier with the third-party sample data;

[0168] Use the model training data and the third-party sample data associated with the user identifier to train the baseline intent recognition model to generate the third-party intent recognition model.

[0169] Optionally, in an embodiment of the present application, the user identifier includes at least one of the brand name, APP name, and product name corresponding to the third-party user.

[0170] Optionally, in an embodiment of the present application, after generating the third-party intent recognition model, it further includes:

[0171] Obtain the speech text corresponding to the speech signal;

[0172] Input the speech text into at least one baseline intent recognition model and at least one third-party intent recognition model respectively. The at least one baseline intent recognition model outputs a first intent set, and the at least one third-party intent recognition model outputs a second intent set;

[0173] Determine the intent of the speech text according to the confidence of the intents in the first intent set and the confidence of the intents in the second intent set.

[0174] The implementation manners of the above various embodiments can refer to the descriptions of the relevant contents in the specification, and will not be elaborated here.

[0175] Corresponding to the above method for determining the user intent, on the other hand, the present application further provides a device for determining the user intent, and the device includes:

[0176] A speech recognition module, configured to obtain a speech text corresponding to a speech signal;

[0177] A dialogue management module, configured to respectively input the speech text into at least one benchmark intent recognition model and at least one third-party intent recognition model, output a first intent set through the at least one benchmark intent recognition model, and output a second intent set through the at least one third-party intent recognition model, where the third-party intent recognition model is trained based on the benchmark intent recognition model of the same skill category and its model training data; and configured to determine the intent of the speech text according to the confidence of the intents in the first intent set and the confidence of the intents in the second intent set.

[0178] Optionally, in an embodiment of the present application, the third-party intent recognition model is trained in the following manner:

[0179] Obtain a benchmark intent recognition model of a preset skill category and the model training data of the benchmark intent recognition model, where the model training data at least includes a plurality of preset intents and the benchmark sample data and benchmark model parameters respectively corresponding to the plurality of preset intents;

[0180] Obtain third-party sample data matching the preset skill category;

[0181] Use the model training data and the third-party sample data to train the benchmark intent recognition model to generate the third-party intent recognition model.

[0182] Optionally, in an embodiment of the present application, the obtaining third-party sample data matching the preset skill category includes:

[0183] Obtain third-party intents added by a third-party user and the third-party sample data corresponding to the third-party intents, where the third-party intents match the preset skill category, or

[0184] Obtain sample data added by a third-party user on the basis of the benchmark sample data corresponding to the preset intent.

[0185] Optionally, in an embodiment of the present application, training the benchmark intent recognition model using the model training data and the third-party sample data to generate the third-party intent recognition model includes:

[0186] Obtain the user identifier of the third-party user;

[0187] Associate the user identifier with the third-party sample data;

[0188] Train the benchmark intent recognition model using the model training data and the third-party sample data associated with the user identifier to generate the third-party intent recognition model.

[0189] Optionally, in an embodiment of the present application, the user identifier includes at least one of the brand name, APP name, and product name corresponding to the third-party user.

[0190] Optionally, in an embodiment of the present application, determining the intent of the speech text according to the confidence levels of the intents in the first intent set and the confidence levels of the intents in the second intent set includes:

[0191] When it is determined that the confidence levels of all the intents included in the first intent set are less than or equal to a first preset threshold and the confidence levels of all the intents included in the second intent set are greater than a second preset threshold, use the intent with the highest confidence level in the second intent set as the intent of the speech text; or,

[0192] When it is determined that the confidence levels of all the intents included in the second intent set are less than or equal to the second preset threshold, use the intent with the highest confidence level in the first intent set as the intent of the speech text; or,

[0193] When it is determined that the confidence levels of all the intents included in the first intent set are greater than or equal to the first preset threshold and the confidence levels of all the intents included in the second intent set are greater than the second preset threshold, use the intent with the highest confidence level in the first intent set and the second intent set as the intent of the speech text.

[0194] Optionally, in an embodiment of the present application, the first preset threshold and the second preset threshold are set to match the corresponding skill categories respectively.

[0195] Corresponding to the above method for generating an intent recognition model, on the other hand, the present application also provides a device for generating an intent recognition model, and the device includes:

[0196] A skill category acquisition module, configured to acquire a preset skill category selected by a third-party user;

[0197] A model acquisition module, configured to acquire a benchmark intent recognition model corresponding to the preset skill category and its model training data;

[0198] A sample acquisition module, configured to acquire third-party sample data from the third-party user that matches the preset skill category;

[0199] A model generation module, configured to train the benchmark intent recognition model using the model training data and the third-party sample data to generate a third-party intent recognition model, where the third-party intent recognition model is an intent recognition model corresponding to the third-party user.

[0200] Optionally, in an embodiment of the present application, acquiring the third-party sample data that matches the preset skill category includes:

[0201] Acquiring third-party intents added by the third-party user and third-party sample data corresponding to the third-party intents, where the third-party intents match the preset skill category, or,

[0202] Acquiring sample data added by the third-party user based on the benchmark sample data corresponding to the preset intent.

[0203] Optionally, in an embodiment of the present application, training the benchmark intent recognition model using the model training data and the third-party sample data to generate the third-party intent recognition model includes:

[0204] Acquiring a user identifier of the third-party user;

[0205] Associating the user identifier with the third-party sample data;

[0206] Training the benchmark intent recognition model using the model training data and the third-party sample data associated with the user identifier to generate the third-party intent recognition model.

[0207] Optionally, in an embodiment of the present application, the user identifier includes at least one of a brand name, an APP name, and a product name corresponding to the third-party user.

[0208] Optionally, in an embodiment of the present application, it further includes:

[0209] A speech recognition module, configured to acquire a speech text corresponding to a speech signal;

[0210] A dialogue management module, configured to input the speech text into at least one baseline intent recognition model and at least one third-party intent recognition model respectively, output a first intent set through the at least one baseline intent recognition model, and output a second intent set through the at least one third-party intent recognition model; and, configured to determine the intent of the speech text according to the confidence levels of the intents in the first intent set and the confidence levels of the intents in the second intent set.

[0211] An embodiment of the present application provides a terminal device, such as Figure 7 as shown, the terminal device includes: a processor and a memory for storing processor-executable instructions; wherein, the processor is configured to execute the instructions so that the terminal device implements the above method.

[0212] An embodiment of the present application provides a non-volatile computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.

[0213] An embodiment of the present application provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, and when the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0214] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded device, such as a punched card or raised structure in a groove storing instructions thereon, and any suitable combination of the foregoing.

[0215] The computer-readable program instructions or code described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0216] The computer program instructions for performing the operations of this application can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., using an Internet service provider to connect through the Internet). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of this application.

[0217] Aspects of the present application are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0218] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0219] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0220] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of apparatus, systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved.

[0221] It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by hardware (such as circuits or an ASIC (Application Specific Integrated Circuit)) that performs the corresponding functions or acts, or can be implemented by a combination of hardware and software, such as firmware.

[0222] Although the present invention has been described in conjunction with various embodiments, it will be understood by those skilled in the art that other variations of the disclosed embodiments can be understood and effected while practicing the claimed invention, by studying the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single processor or other unit may implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not indicate that these measures cannot be combined to advantage.

[0223] The embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A method for determining user intent, characterized in that, it includes: Obtain the speech text corresponding to the speech signal; Input the speech text into at least one benchmark intent recognition model and at least one third-party intent recognition model respectively. The first intent set is output by the at least one benchmark intent recognition model, and the second intent set is output by the at least one third-party intent recognition model. Among them, the third-party intent recognition model is trained based on the benchmark intent recognition model of the same skill category, the model training data of the benchmark intent recognition model, and third-party sample data; Determine the intent of the speech text according to the confidence of the intent in the first intent set and the confidence of the intent in the second intent set.

2. The method according to claim 1, characterized in that, the third-party intent recognition model is trained in the following manner: Obtain the benchmark intent recognition model of the preset skill category and the model training data of the benchmark intent recognition model. The model training data at least includes a plurality of preset intents and the benchmark sample data and benchmark model parameters respectively corresponding to the plurality of preset intents; Obtain third-party sample data matching the preset skill category; Use the model training data and the third-party sample data to train the benchmark intent recognition model to generate the third-party intent recognition model.

3. The method according to claim 2, characterized in that, the obtaining of the third-party sample data matching the preset skill category includes: Obtain the third-party intent added by the third-party user and the third-party sample data corresponding to the third-party intent. The third-party intent matches the preset skill category, or, Obtain the sample data added by the third-party user on the basis of the benchmark sample data corresponding to the preset intent.

4. The method according to claim 2, characterized in that, the using the model training data and the third-party sample data to train the benchmark intent recognition model to generate the third-party intent recognition model includes: Obtain the user identification of the third-party user; Associate the user identification with the third-party sample data; Use the model training data and the third-party sample data associated with the user identification to train the benchmark intent recognition model to generate the third-party intent recognition model.

5. The method according to claim 4, characterized in that, the user identification includes at least one of the brand name, APP name, and product name corresponding to the third-party user.

6. The method according to claim 1, characterized in that, the determining the intent of the speech text according to the confidence of the intent in the first intent set and the confidence of the intent in the second intent set includes: In the case where it is determined that the confidence of the intent included in the first intent set is less than or equal to the first preset threshold, and the confidence of the intent included in the second intent set is greater than the second preset threshold, use the intent with the highest confidence in the second intent set as the intent of the speech text; or, In the case where the confidence levels of the intents included in the second intent set are all less than or equal to a second preset threshold, the intent with the highest confidence level in the first intent set is used as the intent of the speech text; or, In the case where the confidence levels of the intents included in the first intent set are all greater than or equal to a first preset threshold and the confidence levels of the intents included in the second intent set are all greater than a second preset threshold, the intent with the highest confidence level in the first intent set and the second intent set is used as the intent of the speech text.

7. The method according to claim 6, characterized in that, the first preset threshold and the second preset threshold are set to match the corresponding skill categories respectively.

8. A method for generating an intent recognition model, characterized in that, comprising: obtaining a preset skill category selected by a third-party user; obtaining a benchmark intent recognition model corresponding to the preset skill category and its model training data; obtaining third-party sample data from the third-party user that matches the preset skill category; training the benchmark intent recognition model using the model training data and the third-party sample data to generate a third-party intent recognition model, and the third-party intent recognition model is an intent recognition model corresponding to the third-party user.

9. The method according to claim 8, characterized in that, the obtaining of the third-party sample data that matches the preset skill category includes: obtaining third-party intents added by the third-party user and third-party sample data corresponding to the third-party intents, where the third-party intents match the preset skill category, or, obtaining sample data added by the third-party user based on benchmark sample data corresponding to a preset intent.

10. The method according to claim 8, characterized in that, the training of the benchmark intent recognition model using the model training data and the third-party sample data to generate the third-party intent recognition model includes: obtaining a user identifier of the third-party user; associating the user identifier with the third-party sample data; training the benchmark intent recognition model using the model training data and the third-party sample data associated with the user identifier to generate the third-party intent recognition model.

11. The method according to claim 10, characterized in that, the user identifier includes at least one of a brand name, an APP name, and a product name corresponding to the third-party user.

12. The method according to claim 8, characterized in that, after generating the third-party intent recognition model, it further includes: obtaining a speech text corresponding to a speech signal; inputting the speech text into at least one benchmark intent recognition model and at least one third-party intent recognition model respectively, outputting a first intent set through the at least one benchmark intent recognition model, and outputting a second intent set through the at least one third-party intent recognition model; determining the intent of the speech text according to the confidence levels of the intents in the first intent set and the confidence levels of the intents in the second intent set.

13. A device for determining a user intent, characterized in that, comprising: A speech recognition module, configured to obtain the speech text corresponding to the speech signal; A dialogue management module, configured to input the speech text into at least one benchmark intent recognition model and at least one third-party intent recognition model respectively. The first intent set is output by the at least one benchmark intent recognition model, and the second intent set is output by the at least one third-party intent recognition model. Wherein, the third-party intent recognition model is trained based on the benchmark intent recognition model of the same skill category, the model training data of the benchmark intent recognition model, and third-party sample data; And, configured to determine the intent of the speech text according to the confidence of the intent in the first intent set and the confidence of the intent in the second intent set.

14. The apparatus according to claim 13, wherein, the third-party intent recognition model is trained in the following manner: Obtain the benchmark intent recognition model of the preset skill category and the model training data of the benchmark intent recognition model. The model training data at least includes a plurality of preset intents and the benchmark sample data and benchmark model parameters respectively corresponding to the plurality of preset intents; Obtain third-party sample data matching the preset skill category; Use the model training data and the third-party sample data to train the benchmark intent recognition model to generate the third-party intent recognition model.

15. The apparatus according to claim 14, wherein, the obtaining of the third-party sample data matching the preset skill category includes: Obtain the third-party intent added by the third-party user and the third-party sample data corresponding to the third-party intent. The third-party intent matches the preset skill category, or, Obtain the sample data added by the third-party user on the basis of the benchmark sample data corresponding to the preset intent.

16. The apparatus according to claim 14, wherein, the using of the model training data and the third-party sample data to train the benchmark intent recognition model to generate the third-party intent recognition model includes: Obtain the user identifier of the third-party user; Associate the user identifier with the third-party sample data; Use the model training data and the third-party sample data associated with the user identifier to train the benchmark intent recognition model to generate the third-party intent recognition model.

17. The apparatus according to claim 16, wherein, the user identifier includes at least one of the brand name, APP name, and product name corresponding to the third-party user.

18. The apparatus according to claim 13, wherein, the determining of the intent of the speech text according to the confidence of the intent in the first intent set and the confidence of the intent in the second intent set includes: In the case where it is determined that the confidence of the intent included in the first intent set is less than or equal to the first preset threshold, and the confidence of the intent included in the second intent set is greater than the second preset threshold, use the intent with the highest confidence in the second intent set as the intent of the speech text; or, In the case where the confidence levels of the intents included in the second intent set are all less than or equal to a second preset threshold, the intent with the highest confidence level in the first intent set is taken as the intent of the speech text; or, In the case where the confidence levels of the intents included in the first intent set are all greater than or equal to a first preset threshold and the confidence levels of the intents included in the second intent set are all greater than a second preset threshold, the intent with the highest confidence level in the first intent set and the second intent set is taken as the intent of the speech text.

19. The apparatus according to claim 18, wherein, the first preset threshold and the second preset threshold are set to match the corresponding skill categories respectively.

20. An apparatus for generating an intent recognition model, wherein, comprises: a skill category acquisition module, configured to acquire a preset skill category selected by a third-party user; a model acquisition module, configured to acquire a benchmark intent recognition model corresponding to the preset skill category and its model training data; a sample acquisition module, configured to acquire third-party sample data from the third-party user that matches the preset skill category; a model generation module, configured to train the benchmark intent recognition model using the model training data and the third-party sample data to generate a third-party intent recognition model, and the third-party intent recognition model is an intent recognition model corresponding to the third-party user.

21. The apparatus according to claim 20, wherein, the acquiring the third-party sample data that matches the preset skill category includes: acquiring third-party intents added by the third-party user and third-party sample data corresponding to the third-party intents, where the third-party intents match the preset skill category, or, acquiring sample data added by the third-party user based on benchmark sample data corresponding to a preset intent.

22. The apparatus according to claim 20, wherein, the training the benchmark intent recognition model using the model training data and the third-party sample data to generate the third-party intent recognition model includes: acquiring a user identifier of the third-party user; associating the user identifier with the third-party sample data; training the benchmark intent recognition model using the model training data and the third-party sample data associated with the user identifier to generate the third-party intent recognition model.

23. The apparatus according to claim 22, wherein, the user identifier includes at least one of a brand name, an APP name, and a product name corresponding to the third-party user.

24. The apparatus according to claim 20, wherein, further comprises: a speech recognition module, configured to acquire a speech text corresponding to a speech signal; a dialogue management module, configured to input the speech text into at least one benchmark intent recognition model and at least one third-party intent recognition model respectively, and output a first intent set through the at least one benchmark intent recognition model and output a second intent set through the at least one third-party intent recognition model; and for determining the intent of the speech text according to the confidence of the intents in the first intent set and the confidence of the intents in the second intent set.

25. A terminal device, characterized in that it includes: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions such that the terminal device implements the method according to any one of claims 1-7, or implements the method according to any one of claims 8-12.

26. A non-volatile computer-readable storage medium, on which computer program instructions are stored, characterized in that when the computer program instructions are executed by a processor, the method according to any one of claims 1-7 is implemented, or the method according to any one of claims 8-12 is implemented.

27. A computer program product, characterized in that it includes computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, and when the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes to implement the method according to any one of claims 1-7, or implements the method according to any one of claims 8-12.

Citation Information

Patent Citations

  • Intention identification method and device based on text classification, equipment and storage medium

    CN110147445A

  • Determining domains for natural language understanding

    US10453117B1