User intention recognition method, device and electronic device
By setting the intent value in the intent list to generate a personalized intent list, the server determines the user's intent based on the voice data and intent values, solving the problem of low recognition rate of the TV in intelligent voice interaction scenarios, and improving the accuracy of the user's intent recognition.
Patent Information
- Application Number
- CN202111660073.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2041-12-30
AI Technical Summary
In the prior art, the TV set has a low recognition rate of user intentions in intelligent voice interaction scenarios, and often requires the user to enter voice multiple times to respond.
The user sets the intent value for each intent service in the intent list and generates a personalized intent list. The server determines the user's intent based on the voice data and intent values, and retrieves the corresponding service data and sends it to the device.
The recognition rate of user intentions is increased, allowing the server to determine user intentions more accurately, and reducing the number of times the user inputs is performed.
Smart Images

Figure CN114187897B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a method, device, and electronic device for identifying user intention. Background Art
[0002] With the development of voice recognition technology, intelligent voice interaction technology has gradually become a standard feature of terminal devices (such as mobile phones, tablets, or smart home products such as smart appliances). In intelligent voice interaction scenarios, users can control smart home appliances through voice. Taking a television as an example, users can control the TV through voice to perform a series of TV control operations such as watching videos, listening to music, or checking the weather. However, in actual operation, the TV often fails to correctly understand the user's intention due to unclear or ambiguous voice input. The user needs to input voice multiple times before it can respond, resulting in a low recognition rate of the actual user's intention. Summary of the Invention
[0003] The present application provides a method, device and electronic device for identifying user intention, which solves the problem of low recognition rate of user intention by television sets in the prior art.
[0004] To achieve the above objectives, this application adopts the following technical solutions:
[0005] In a first aspect, the present application provides a method for identifying user intention, comprising: receiving voice data sent by a first device; wherein the account currently logged in to the first device is the first account; obtaining a personalized intention list corresponding to the first account; wherein the personalized intention list includes intention values corresponding to different intention services set by the user; determining the user intention based on the voice text and intention value corresponding to the voice data; wherein the user intention is any one of the different intention services; calling the service data containing the voice text in the user intention; and sending the service data to the first device.
[0006] In some feasible examples, the user intent is determined based on the voice text and intent value corresponding to the voice data, including: determining the first confidence corresponding to each intention service based on the voice text and intent value corresponding to the voice data; and determining that the intention service with the first confidence greater than or equal to the first confidence threshold is the user intention.
[0007] In some feasible examples, the method for identifying user intent provided in the present application also includes: determining the second confidence corresponding to each intent business based on the first confidence and the intent value when determining that there is no intent business with a first confidence greater than or equal to a first confidence threshold; and determining that the intent business with a second confidence greater than or equal to the second confidence threshold is the user intent.
[0008] In some feasible examples, the method for identifying user intent provided in the present application further includes: determining that the default intent service is the user intent when it is determined that there is no intent service with a second confidence level greater than or equal to a second confidence threshold.
[0009] In some feasible examples, the first confidence level corresponding to each intent service is determined based on the voice text and intent value corresponding to the voice data, including: inputting the voice text and intent value corresponding to the voice data into a pre-configured intent recognition model to determine the first confidence level corresponding to each intent service.
[0010] In some feasible examples, before receiving the voice data sent by the first device, the user intent recognition method provided in this application also includes: receiving an intent setting request sent by the first device; in response to the intent setting request, if it is determined that the personalized intent list corresponding to the first account is not saved, sending a default intent list to the first device; receiving the personalized intent list sent by the first device; establishing a correspondence between the first account and the personalized intent list, and saving the personalized intent list.
[0011] In a second aspect, the present application provides a method for identifying user intentions, which is applied to a first device, including: sending voice data to a server; wherein the account currently logged in to the first device is the first account, and the first account corresponds to a personalized intent list, and the personalized intent list includes intention values corresponding to different intent services set by the user; receiving business data sent by the server; wherein the business data includes business data containing voice text in the user intention, and the user intention is determined based on the voice text and intention value corresponding to the voice data, and the user intention is any one of the different intent services.
[0012] In some feasible examples, the method for identifying user intent provided in the present application also includes: sending an intent setting request to a server; receiving a default intent list sent by the server; and sending a personalized intent list to the server in response to the user's setting operation on the default intent list.
[0013] On the third aspect, the present application provides a device for identifying user intention, including: a transceiver unit for receiving voice data sent by a first device; wherein the account currently logged in to the first device is the first account; the transceiver unit is also used to obtain a personalized intention list corresponding to the first account; wherein the personalized intention list includes intention values corresponding to different intention services set by the user; a processing unit is used to determine the user intention based on the voice text corresponding to the voice data received by the transceiver unit and the intention value obtained by the transceiver unit; wherein the user intention is any one of the different intention services; the processing unit is also used to call the service data in the user intention that contains the voice text received by the transceiver unit; the processing unit is also used to control the transceiver unit to send the service data to the first device.
[0014] In some feasible examples, the processing unit is specifically used to determine the first confidence corresponding to each intention service based on the voice text corresponding to the voice data received by the transceiver unit and the intention value obtained by the transceiver unit; the processing unit is specifically used to determine that the intention service with the first confidence greater than or equal to the first confidence threshold is the user intention.
[0015] In some feasible examples, the processing unit is further used to determine the second confidence corresponding to each intended service based on the first confidence and the intention value obtained by the transceiver unit when there is no intended service with a first confidence greater than or equal to the first confidence threshold; the processing unit is also used to determine that the intended service with a second confidence greater than or equal to the second confidence threshold is the user intention.
[0016] In some feasible examples, the processing unit is further configured to determine that the default intended service is the user intention when it is determined that there is no intended service with a second confidence level greater than or equal to a second confidence level threshold.
[0017] In some feasible examples, the processing unit is specifically used to input the voice text corresponding to the voice data received by the transceiver unit and the intention value obtained by the transceiver unit into a pre-configured intention recognition model to determine the first confidence corresponding to each intention service.
[0018] In some feasible examples, the transceiver unit is further used to receive an intent setting request sent by the first device; the processing unit is further used to, in response to the intent setting request received by the transceiver unit, control the transceiver unit to send a default intent list to the first device when it is determined that the personalized intent list corresponding to the first account has not been saved; the transceiver unit is further used to receive the personalized intent list sent by the first device; the processing unit is further used to establish a correspondence between the first account and the personalized intent list received by the transceiver unit, and save the personalized intent list.
[0019] In the fourth aspect, the present application provides a user intention recognition device including: a transceiver unit for sending voice data to a server; wherein, the account currently logged in by the first device is the first account, and the first account corresponds to a personalized intent list, and the personalized intent list includes intention values corresponding to different intent services set by the user; the transceiver unit is also used to receive business data sent by the server; wherein, the business data includes business data containing voice text in the user's intention, and the user's intention is determined based on the voice text and intention value corresponding to the voice data, and the user's intention is any one of the different intent services.
[0020] In some feasible examples, the identification device also includes a processing unit; a transceiver unit, further used to send an intent setting request to a server; the transceiver unit, further used to receive a default intent list sent by the server; and the processing unit, further used to control the transceiver unit to send a personalized intent list to the server in response to a user's setting operation on the default intent list received by the transceiver unit.
[0021] In a fifth aspect, the present application provides a speech recognition system, comprising any server provided in the third aspect, and any electronic device provided in the fourth aspect.
[0022] In a sixth aspect, the present application provides a computer-readable storage medium comprising instructions, which, when executed on a computer, enables the computer to execute the method for identifying user intention as described in any one of the items provided in the first aspect.
[0023] In the seventh aspect, the present application provides a server, comprising: a communication interface, a processor, a memory, and a bus; the memory is used to store computer-executable instructions, and the processor is connected to the memory through the bus; when the server is running, the processor executes the computer-executable instructions stored in the memory, so that the server executes the user intention recognition method as described in any one of the items provided in the first aspect.
[0024] In an eighth aspect, the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method for identifying user intention as described in the design method of the first aspect.
[0025] In the ninth aspect, the present application provides a computer-readable storage medium comprising instructions, which, when executed on a computer, enables the computer to execute the method for identifying user intention as described in any one of the items provided in the second aspect.
[0026] In the tenth aspect, the present application provides an electronic device, comprising: a communication interface, a processor, a memory, and a bus; the memory is used to store computer-executable instructions, and the processor is connected to the memory through the bus; when the electronic device is running, the processor executes the computer-executable instructions stored in the memory, so that the electronic device performs the user intention recognition method as described in any one of the items provided in the second aspect.
[0027] In the eleventh aspect, the present application provides a computer program product, which, when running on a computer, enables the computer to execute the method for identifying user intention as described in the design method of the second aspect.
[0028] It should be noted that the above-mentioned computer instructions may be stored in whole or in part on a first computer-readable storage medium. The first computer-readable storage medium may be packaged together with the server, or may be packaged separately with the electronic device or the server processor, and this application does not limit this.
[0029] The descriptions of the third, sixth, seventh and eighth aspects of this application can refer to the detailed description of the first aspect; and the beneficial effects of the descriptions of the third, sixth, seventh and eighth aspects can refer to the analysis of the beneficial effects of the first aspect, which will not be repeated here.
[0030] The descriptions of the fourth, ninth, tenth and eleventh aspects of this application can refer to the detailed description of the second aspect; and the beneficial effects of the descriptions of the fourth, ninth, tenth and eleventh aspects can refer to the beneficial effect analysis of the second aspect, which will not be repeated here.
[0031] In this application, the names of the aforementioned servers or electronic devices do not limit the devices or functional modules themselves. In actual implementation, these devices or functional modules may appear with other names. As long as the functions of each device or functional module are similar to those of this application, they fall within the scope of the claims of this application and their equivalents.
[0032] These and other aspects of the present application will become more readily apparent from the following description.
[0033] The present application provides a method for identifying user intentions, whereby the user sets an intention value for each intention service in the intention list, thereby obtaining a personalized intention list for the user. In this way, when the user uses the first device, the first device can send the user's voice data to the server, so that the server can determine the user's personalized intention list based on the first account logged in by the user on the first device. Furthermore, the server can determine the user intention of the user based on the voice text corresponding to the voice data and the intention value corresponding to each intention service. Afterwards, the server retrieves the service data containing the voice text in the user's intention and sends the service data to the first device. Since the user has set the intention value for each intention service on the first device in advance, when the server determines the user intention of the user, it can more accurately determine the user intention based on the intention value for each intention service set by the user, thereby improving the recognition rate of the user intention. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 One of the scenario diagrams of the method for identifying user intention provided in an embodiment of the present application;
[0035] Figure 2 This is one of the structural schematic diagrams of the display device in the method for identifying user intention provided in an embodiment of the present application;
[0036] Figure 3 This is a second structural diagram of a display device in the method for identifying user intention provided in an embodiment of the present application;
[0037] Figure 4 This is a flowchart of a method for identifying user intent provided in an embodiment of the present application;
[0038] Figure 5 This is a second scenario diagram of the method for identifying user intent provided in an embodiment of the present application;
[0039] Figure 6 This is a second flow chart of the method for identifying user intent provided in an embodiment of the present application;
[0040] Figure 7 The third flowchart of the method for identifying user intent provided in an embodiment of the present application;
[0041] Figure 8 A schematic diagram of the structure of the server provided in the embodiment of the present application;
[0042] Figure 9 One of the schematic diagrams of a chip system provided in an embodiment of the present application;
[0043] Figure 10 A schematic diagram of the structure of a television provided in an embodiment of the present application;
[0044] Figure 11 This is a second schematic diagram of a chip system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to make the purpose, implementation mode and advantages of the present application clearer, the exemplary implementation mode of the present application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only part of the embodiments of the present application, not all of the embodiments.
[0046] Based on the exemplary embodiments described in this application, all other embodiments obtained by those of ordinary skill in the art without making any creative work shall fall within the scope of protection of the claims attached to this application. In addition, although the disclosure in this application is introduced according to one or several exemplary examples, it should be understood that each aspect of these disclosures can also constitute a complete embodiment separately. It should be noted that the brief description of the terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.
[0047] Currently, with the rapid development of artificial intelligence (AI), intelligent voice interaction technology is gradually becoming a standard feature of terminal devices (such as mobile phones, tablets, or smart home products such as smart appliances). In intelligent voice interaction scenarios, the most critical aspect of human-computer dialogue technology is the recognition of user intent, that is, the recognition of the intention expressed by the user's voice input data. However, in actual operation, taking the terminal device as a television as an example, the television often fails to correctly understand the user's intention because the user's voice input is unclear or ambiguous, requiring the user to input voice multiple times before responding, resulting in a low recognition rate of the actual user's intention.
[0048] In order to solve the above problems, an embodiment of the present application provides a method for identifying user intentions, whereby the user sets an intention value for each intention service in the intention list, thereby obtaining a personalized intention list for the user. In this way, when the user uses the first device, the first device can send the user's voice data to the server, so that the server can determine the user's personalized intention list based on the first account logged in by the user on the first device. Furthermore, the server can determine the user intention of the user based on the voice text corresponding to the voice data and the intention value corresponding to each intention service. Afterwards, the server retrieves the service data containing the voice text in the user's intention and sends the service data to the first device. Since the user has set the intention value for each intention service on the first device in advance, when the server determines the user intention of the user, it can more accurately determine the user intention based on the intention value for each intention service set by the user, thereby improving the recognition rate of the user intention.
[0049] Figure 1 is a schematic diagram of an operation scenario between a display device and a control device according to one or more embodiments of the present application, such as Figure 1 As shown, a user can operate the display device 200 via a mobile terminal 300 and a control device 100. The control device 100 can be a remote controller, and communication between the remote controller and the display device includes infrared protocol communication, Bluetooth protocol communication, wireless or other wired methods to control the display device 200. The user can control the display device 200 by inputting user commands through buttons on the remote controller, voice input, control panel input, etc. In some embodiments, a mobile terminal, tablet computer, computer, laptop computer, and other smart devices can also be used to control the display device 200.
[0050] In some embodiments, the mobile terminal 300 can install software applications with the display device 200, and achieve connection and communication through a network communication protocol to achieve the purpose of one-to-one control operation and data communication. The audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 to achieve a synchronous display function. The display device 200 also communicates data with the server 400 through a variety of communication methods. The display device 200 can be allowed to communicate and connect through a local area network (LAN), a wireless local area network (WLAN) and other networks. The server 400 can provide various content and interactions to the display device 200. The display device 200 can be a liquid crystal display, an OLED display, or a projection display device. In addition to providing a broadcast receiving television function, the display device 200 can also provide an intelligent network TV function that provides computer support functions.
[0051] In some embodiments, the first device provided by the embodiment of the present application may be the above-mentioned display device 200. The display device 200 is used to send the user's voice data to the server 400, so that the server 400 can determine the user's personalized intention list based on the first account logged in by the user on the display device 200. Furthermore, the server 400 can determine the user's user intention based on the voice text corresponding to the voice data and the intention value corresponding to each intention service. Afterwards, the server 400 retrieves the service data containing the voice text in the user's intention and sends the service data to the display device 200.
[0052] Figure 2 FIG. 2 shows a hardware configuration block diagram of a display device 200 according to an exemplary embodiment. Figure 2 The display device 200 shown includes at least one of a tuner and demodulator 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface 280. The controller includes a central processing unit (CPU), a video processor, an audio processor, a graphics processor, RAM, ROM, and first to nth interfaces for input / output. The display 260 can be at least one of a liquid crystal display (LCD), an OLED display, a touch display, and a projection display, and can also be a projection device and projection screen. The tuner and demodulator 210 receives broadcast television signals via wired or wireless reception, and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals. The detector 230 is used to collect signals from the external environment or external interactions. The controller 250 and tuner and demodulator 210 can be located in different separate devices, that is, the tuner and demodulator 210 can also be located in a device external to the main device where the controller 250 is located, such as an external set-top box.
[0053] In some embodiments, the controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200. The user can enter user commands through the graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input commands through the graphical user interface (GUI). Alternatively, the user can enter user commands by inputting specific sounds or gestures, and the user input interface receives the user input commands by recognizing the sounds or gestures through sensors.
[0054] In some embodiments, the sound collector can be a microphone, also known as a "microphone" or "microphone", which is used to convert sound signals into electrical signals. When performing voice interaction, the user can put their mouth close to the microphone to speak and input the sound signal into the microphone. The display device 200 can be provided with at least one microphone. In other embodiments, the display device 200 can be provided with two microphones, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the display device 200 can also be provided with three, four or more microphones to realize sound signal collection, noise reduction, and identification of sound sources, and realize directional recording function, etc.
[0055] Among them, the microphone may be built into the display device 200, or the microphone may be connected to the display device 200 by wire or wireless means. For example, the microphone may be arranged at the lower edge of the display 260 of the display device 200. Of course, the embodiment of the present application does not limit the position of the microphone on the display device 200. Alternatively, the display device 200 may not include a microphone, that is, the above-mentioned microphone is not arranged in the display device 200. The display device 200 may be connected to an external microphone (also referred to as a microphone) through an interface (such as a USB interface 130). The external microphone may be fixed to the display device 200 by an external fixing member (such as a camera bracket with a clip). For example, the external microphone may be fixed to the edge of the display 260 of the display device 200, such as the upper edge, by an external fixing member.
[0056] In some embodiments, a "user interface" is a medium interface for interaction and information exchange between an application or operating system and a user. It enables the conversion between the internal form of information and a form acceptable to the user. A common form of user interface is a graphical user interface (GUI), which refers to a user interface related to computer operations that is displayed graphically. It can be an interface element such as an icon, window, or control displayed on the display of an electronic device. A control can include at least one of the following visual interface elements: an icon, button, menu, tab, text box, dialog box, status bar, navigation bar, widget, etc.
[0057] In some examples, the display device 200 of one or more embodiments is a TV set 1, and the operating system of the TV set 1 is an Android system. Figure 3 As shown, the television 1 can be logically divided into an application layer (abbreviated as “application layer”) 21 , a kernel layer 22 and a hardware layer 23 .
[0058] Among them, such as Figure 3 As shown, the hardware layer may include Figure 2The controller 250, communicator 220, detector 230, and display 260 are shown. The application layer 21 includes one or more applications. The applications can be system applications or third-party applications. For example, the application layer 21 includes a voice recognition application that can provide a voice interaction interface and services for connecting the television 1 to the server 400.
[0059] The kernel layer 22 serves as a software middleware between the hardware layer and the application layer 21 and is used to manage and control hardware and software resources.
[0060] Server 400 includes a communication control module 201, a semantic central control module 202, an intent recognition module 203, a business system module 204, and a data storage module 205. Communication control module 201 is used to establish a communication connection with television 1. For example, the speech recognition application in television 1 establishes a communication connection with communication control module 201 of server 400 by calling communicator 220.
[0061] In some examples, the kernel layer 22 includes a detector driver, which is configured to send voice data collected by the detector 230 to a voice recognition application. When the voice recognition application in the television 1 is activated and a communication connection is established between the television 1 and the server 400, the detector driver is configured to send the user input voice data collected by the detector 230 to the voice recognition application. The voice recognition application then sends the voice data to the semantic central control module 202 of the server. After receiving the voice data from the television 1, the semantic central control module 202 determines the voice text corresponding to the voice data. The semantic central control module 202 sends a request to the data storage module 205 to query the personalized intent list corresponding to the first account currently logged into the television 1. Upon receiving the request from the semantic central control module 202, the data storage module queries the personalized intent list corresponding to the first account currently logged into the television 1. The data storage module 205 then sends the personalized intent list to the semantic central control module 202. The semantic central control module 202 sends the voice text and intent value corresponding to the voice data to the intent recognition module 203. The intention recognition module 203 determines the user intention based on the voice text corresponding to the voice data sent by the semantic central control module 202 and the intention value sent by the semantic central control module 202. The intention recognition module 203 sends the determined user intention to the semantic central control module 202. The semantic central control module 202 sends a business request to the business system module 204 to call the business data containing the voice text in the user intention. After receiving the business request sent by the semantic central control module 202, the business system module 204 calls the business data containing the voice text in the user intention. The business system module 204 sends the business data containing the voice text in the user intention to the voice recognition application of the TV 1. After receiving the business data sent by the server 400, the voice recognition application controls the display 192 to display the business data.
[0062] The voice data involved in this application may be data authorized by the user or fully authorized by all parties.
[0063] The methods in the following embodiments can all be implemented in the television set 1 having the above hardware structure. In the following embodiments, the methods of the embodiments of the present application are described by taking the display device 200 as the television set 1 as an example.
[0064] The present application embodiment provides a method for identifying user intentions, such as Figure 4 As shown, the method for identifying the user intention may include S11-S15.
[0065] S11. The server 400 receives the voice data sent by the TV 1. The account currently logged into the TV 1 is the first account.
[0066] S12. The server 400 obtains a personalized intent list corresponding to the first account, wherein the personalized intent list includes intent values corresponding to different intent services set by the user.
[0067] In some examples, a default intent list is pre-stored in the TV 1. The intent list includes intent services and intent default values. The intent default value can be pre-set or determined based on the number of requests for the intent service. For example, the intent default value is set by the operation and maintenance personnel when the TV 1 leaves the factory. Alternatively, after collecting the number of requests for intent services of a preset duration, the server 1 sorts the number of requests in descending order to determine the sorting of the intent services. Finally, the intent default value of each intent service is determined based on the ratio of the number of requests for each intent service to the total number of requests.
[0068] For example, the intent service includes five user intents, namely, opening an application, searching for videos, playing music, checking the weather, and singing karaoke. The default value of the intent is a star rating, which includes six star ratings, ranging from 0 to 5. The method for obtaining a personalized intent list is as follows:
[0069] Specifically, a higher star rating indicates a higher user interest in the service.
[0070] For example, the default intent list is shown in Table 1.
[0071] Table 1
[0072] Intent Business Star rating Video Search 3 Open the app 3 Music playback 3 Weather query 3 Karaoke 3
[0073] In other examples, when the server 400 cannot identify the user's intention, it can display the following information: Figure 5 The interface 400 shown in (a) of FIG. 400 includes a button 4001 for prompting the user to set the intention. For example, when the server 400 determines that the user intention cannot be recognized for three times in total, the server 400 may display the following information: Figure 5 Alternatively, the server 400 may display the following interface each time it fails to recognize the user's intention: Figure 5 When the user first enters the intention setting interface, the TV 1 displays the following after receiving the user's selection operation on button 4001. Figure 5 Interface 401 shown in (b) of FIG. Interface 401 includes an intention setting interface title bar 4010 and a default intention list 4011. Afterwards, the user can set the required star rating for each intention service, so that the server 400 can determine the user's intention based on the personalized intention list set by the user. After the user sets the star rating corresponding to each intention service, the TV 1 displays the following Figure 5Interface 402 shown in (c) of FIG. Interface 402 includes an intention setting interface title bar 4010, a reset intention list (also called a personalized intention list) 4020, a confirmation save button 4021, and a redo button 4022. After receiving the user's selection operation on the confirmation save button 4021, the TV set 1 sends the personalized intention list to the server 400. After receiving the user's selection operation on the redo button 4022, the TV set 1 displays the following Figure 5 Interface 401 shown in (b).
[0074] For example, the personalized intent list is shown in Table 2.
[0075] Table 2
[0076] Intent Business Star rating Video Search 5 Open the app 3 Music playback 2 Weather query 1 Karaoke 0
[0077] Specifically, the user can delete the intended services in the default intent list, set the star rating of the intended services, and restore the star rating of the set intended services to the default value of the intent. For example, when the user no longer needs karaoke, the karaoke in Table 1 can be deleted, and the updated personalized intent list is shown in Table 3.
[0078] Table 3
[0079] Intent Business Star rating Video Search 3 Open the app 3 Music playback 3 Weather query 3
[0080] like Figure 5 As shown in (c), the user sets the star rating of the video search from star 3 to star 5. The user sets the star rating of the music playback from star 3 to star 2. The user sets the star rating of the weather query from star 3 to star 1. The user sets the star rating of the karaoke from star 3 to star 0. Later, if the user wants to restore the star rating of the set intent service to the default value of the intent, he can do so by Figure 5 Select the redo button 4022 in (c) to restore the star rating of each intended service to the intended default value.
[0081] The above example is based on Figure 5 The selection operation of the reset button 4022 in (c) is explained by taking the star rating of each intended service as an example to restore the intended default value. In other examples, the user can restore the star rating of the intended service to the intended default value, which is not limited here.
[0082] Specifically, in some examples, the intended services in the intent list are sorted according to the size of the star rating, such as arranging the intended services in the intent list in order of star rating from large to small.
[0083] S13: The server 400 determines the user intention based on the voice text and the intention value corresponding to the voice data, wherein the user intention is any one of different intention services.
[0084] In some examples, the server 400 performs text conversion and word segmentation on the voice data sent by the TV 1 to determine the voice text corresponding to the voice data.
[0085] S14. The server 400 calls the service data containing the voice text in the user's intention.
[0086] For example, referring to the example of S12 above, assume that the voice text corresponding to the voice data is "Chinese Paladin" and the user intends to search for videos. In this case, server 400 searches for service data related to "Chinese Paladin" in the video search. Server 1 then sends the found video data related to "Chinese Paladin" to television set 1. In this way, the user can watch the video data related to "Chinese Paladin".
[0087] S15 . The server 400 sends service data to the TV 1 .
[0088] An embodiment of the present application provides a method for identifying user intentions, whereby a user sets an intention value for each intention service in an intention list, thereby obtaining a personalized intention list for the user. In this way, when a user uses a first device, the first device can send the user's voice data to the server, so that the server can determine the user's personalized intention list based on the first account logged in by the user on the first device. Furthermore, the server can determine the user intention of the user based on the voice text corresponding to the voice data and the intention value corresponding to each intention service. Afterwards, the server retrieves the service data containing the voice text in the user's intention and sends the service data to the first device. Since the user has set the intention value for each intention service on the first device in advance, when the server determines the user intention of the user, it can more accurately determine the user intention of the user based on the intention value for each intention service set by the user, thereby improving the recognition rate of the user intention.
[0089] In some examples, combined Figure 4 ,like Figure 6 As shown, the above S13 can be specifically implemented through the following S130 and S131.
[0090] S130. The server 400 determines a first confidence level corresponding to each intended service based on the voice text and the intention value corresponding to the voice data.
[0091] Specifically, a higher first confidence level indicates a higher frequency of user use of the intended service.
[0092] Specifically, the sum of the first confidences corresponding to each intended service is equal to 1.
[0093] S131. The server 400 determines that an intended service with a first confidence level greater than or equal to a first confidence level threshold is a user intention.
[0094] Specifically, the first confidence threshold and the second confidence threshold may be the same or different. In some examples, the first confidence threshold and the second confidence threshold are the same, and both the first confidence threshold and the second confidence threshold are 0.7.
[0095] In some examples, combined Figure 6 ,like Figure 7 As shown, the above S130 can be specifically implemented through the following S1300.
[0096] S1300 , the server 400 inputs the speech text and the intent value corresponding to the speech data into a pre-configured intent recognition model, and determines a first confidence level corresponding to each intent service.
[0097] In some examples, the intent recognition model training process is as follows:
[0098] S1. The server 400 obtains a training sample speech and a labeling result of the training sample speech, wherein the training sample speech includes speech text and user intention.
[0099] S2. The server 400 inputs the training sample speech into the deep learning model.
[0100] Specifically, the deep learning model can be a text convolutional neural network (TEXTCNN).
[0101] S3. The server 400 determines whether the prediction comparison result of the training sample speech output by the deep learning model matches the annotation result based on the target loss function.
[0102] In some examples, the objective loss function is a cross loss function, and the model is optimized by minimizing the loss function, which is generally:
[0103] loss=-∑ i y′log(f(x)).
[0104] Where y′ is the user intention of the input voice text, x is the input voice text, and f is the trained model. The intent recognition model f is obtained by minimizing the loss function.
[0105] In order to improve the recognition rate of user intent, the device method for user intent provided in the embodiment of the present application can incorporate the user's intention preference degree ∝ into the loss function and change the objective function to:
[0106] loss=-∑ i ∝ i y′log(f(x)).
[0107] In this way, the input of the intent recognition model f includes not only the voice text x but also the user's intent preference value ∝. In the device method for user intent provided in the embodiments of the present application, because the training of the intent recognition model f includes guidance from the user's intent preference score, the trained intent recognition model f is more consistent with the user's intent preference, thereby improving the recognition rate of user intent without collecting user privacy data.
[0108] Where i represents the total number of services included in the intended service. As shown in Table 1, the intended services include opening applications, video search, music playback, weather query, and karaoke. In this case, i is equal to 5. i represents the star rating of the i-th intended service, as shown in Table 1. Assuming that i is equal to 1 and the first intended service is video search, then ∝1=3.
[0109] S4. When the predicted comparison result does not match the annotation result, the server 400 iteratively updates the network parameters of the deep learning model repeatedly until the model converges to obtain an intent recognition model.
[0110] In some examples, combined Figure 6 ,like Figure 7 As shown, the method for identifying user intention provided in the embodiment of the present application also includes: S132 and S133.
[0111] S132: When the server 400 determines that there is no intended service with a first confidence greater than or equal to a first confidence threshold, the server 400 determines a second confidence corresponding to each intended service according to the first confidence and the intention value.
[0112] Specifically, the sum of the second confidences corresponding to each intended service is equal to 1.
[0113] In some examples, if the server 400 determines that there is no intended service with a first confidence level greater than or equal to the first confidence level threshold, it means that the server 400 cannot identify the user's intention. In this case, in order to improve the recognition rate of the server 400 in identifying the user's intention, the user intention recognition method provided in the embodiment of the present application needs to determine the second confidence level corresponding to each intended service based on the first confidence level and the intention value. Among them, the larger the second confidence level, the higher the probability that the user will access the intended service.
[0114] Among them, the second confidence
[0115] Among them, p i Indicates the first confidence of the i-th intended service.
[0116] S133. The server 400 determines that the intended service with a second confidence level greater than or equal to a second confidence level threshold is the user intention.
[0117] In some examples, when the intent recognition model f does not consider the user's intent preference value ∝, combined with the example given in S1300 above, the first confidence level corresponding to each intent service is shown in Table 4.
[0118] Table 4
[0119] Intent Business First confidence Video Search 0.4 Open the app 0.4 Music playback 0.2 Weather query 0 Karaoke 0
[0120] It can be seen that when the intent recognition model f does not consider the user's intent preference value ∝, the server 400 determines that there is no intent service with a first confidence greater than or equal to the first confidence threshold.
[0121] In comparison, combined with the example given in S1300 above, when the intent recognition model f considers the user's intent preference value ∝, the first confidence level corresponding to each intent service is shown in Table 5.
[0122] Table 5
[0123] Intent Business First confidence Video Search 0.6 Open the app 0.3 Music playback 0.1 Weather query 0 Karaoke 0
[0124] As can be seen, because intent recognition model f considers the user's intent preference value ∝, the output value of the loss function decreases, thereby improving the first confidence level to a certain extent. For example, the first confidence level of a video search increases from 0.4 to 0.6. Although intent recognition model f considers the user's intent preference value ∝, server 400 determines that there are no intent services with a first confidence level greater than or equal to the first confidence threshold.
[0125] Combined with the above examples, it can be seen that the user intention cannot be well identified based on the intention recognition model alone. In order to improve the recognition rate of the user intention by the server 400, the user intention recognition method provided in the embodiment of the present application needs to determine the second confidence corresponding to each intention business based on the first confidence and the intention value.
[0126] For example, in combination with Table 2 and Table 5, the second confidence level corresponding to each intended service is shown in Table 6.
[0127] Table 6
[0128] Intent Business Second confidence level Video Search 0.732 Open the app 0.220 Music playback 0.048 Weather query 0 Karaoke 0
[0129] As can be seen, since the second confidence level of 0.732 for the video search is greater than the second confidence threshold of 0.7, it can be determined that the user's intent is a video search. In this case, if the voice text is "Chinese Paladin," server 400 will search for business data related to "Chinese Paladin" in the video search.
[0130] In some examples, combined Figure 6 ,like Figure 7 As shown, the method for identifying user intention provided in the embodiment of the present application also includes: S134.
[0131] S134: When the server 400 determines that there is no intended service with a second confidence level greater than or equal to a second confidence level threshold, the server 400 determines that the default intended service is the user intention.
[0132] In some examples, if server 400 determines that there is no intended service with a second confidence level greater than or equal to the second confidence threshold, it indicates that server 400 cannot recognize the user's intent. In this case, server 400 determines the default intended service as the user's intent. For example, if the default intended service may be music playback and the voice text is "Chinese Paladin," server 400 will search for service data related to "Chinese Paladin" in the music playback.
[0133] In some examples, combined Figure 4 ,like Figure 7 As shown, the method for identifying user intention provided in the embodiment of the present application also includes: S16-S19.
[0134] S16 . The server 400 receives the intention setting request sent by the TV 1 .
[0135] In some examples, combined with the example given in Example S12 above, the user can set the star rating corresponding to each intent service by himself. The user can select the intent setting button in the TV 1. After the TV 1 determines that it has received the user's selection operation on the intent setting button, it sends an intent setting request to the server 400. After receiving the intent setting request sent by the TV 1, the server 400 determines that the personalized intent list corresponding to the first account is not saved, and sends a default intent list to the TV 1. After receiving the default intent list, the TV 1 displays the following Figure 5 Interface 401 shown in (b).
[0136] S17 . In response to the intent setting request, if the server 400 determines that the personalized intent list corresponding to the first account is not saved, the server 400 sends a default intent list to the TV 1 .
[0137] S18 . The server 400 receives the personalized intent list sent by the TV 1 .
[0138] S19. The server 400 establishes a correspondence between the first account and the personalized intent list, and saves the personalized intent list.
[0139] In some examples, to facilitate management of the personalized intent list for each account, server 400 needs to establish a correspondence between the first account and the personalized intent list and store the personalized intent list in a database of server 400. Subsequently, server 400 can query the personalized intent list corresponding to each first account based on the established correspondence between the first account and the personalized intent list.
[0140] The present application embodiment provides a method for identifying user intentions, such as Figure 6 As shown, the method for identifying the user intention may include S20 and S21.
[0141] S20: TV 1 sends voice data to server 400. The account currently logged into TV 1 is the first account, which corresponds to a personalized intent list. The personalized intent list includes intent values corresponding to different intent services set by the user.
[0142] S21. The TV 1 receives service data sent by the server 400. The service data includes service data including voice text in the user's intention. The user's intention is determined based on the voice text and the intention value corresponding to the voice data. The user's intention is any one of different intention services.
[0143] In some examples, combined Figure 6 ,like Figure 7 As shown, the method for identifying user intention provided in the embodiment of the present application also includes: S22-S24.
[0144] S22 : The TV 1 sends an intention setting request to the server 400 .
[0145] S23 . The TV 1 receives the default intent list sent by the server 400 .
[0146] S24 . The TV 1 sends the personalized intent list to the server 400 in response to the user's setting operation on the default intent list.
[0147] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of method. In order to realize the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily appreciate that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0148] In the embodiment of the present application, the server and the electronic device can be divided into functional modules according to the above method examples. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical function division. In actual implementation, there may be other division methods.
[0149] like Figure 8 As shown, an embodiment of the present application provides a structural diagram of a server 400. The server 400 includes a transceiver unit 101 and a processing unit 102.
[0150] The transceiver unit 101 is used to receive voice data sent by the first device; wherein, the account currently logged in to the first device is the first account; the transceiver unit 101 is also used to obtain a personalized intent list corresponding to the first account; wherein, the personalized intent list includes intent values corresponding to different intent services set by the user; the processing unit 102 is used to determine the user intent based on the voice text corresponding to the voice data received by the transceiver unit 101 and the intent value obtained by the transceiver unit 101; wherein the user intent is any one of the different intent services; the processing unit 102 is also used to call the service data of the user intent that contains the voice text received by the transceiver unit 101; the processing unit 102 is also used to control the transceiver unit 101 to send service data to the first device.
[0151] In some feasible examples, the processing unit 102 is specifically used to determine the first confidence corresponding to each intention service based on the voice text corresponding to the voice data received by the transceiver unit 101 and the intention value obtained by the transceiver unit 101; the processing unit 102 is specifically used to determine that the intention service with the first confidence greater than or equal to the first confidence threshold is the user intention.
[0152] In some feasible examples, the processing unit 102 is further used to determine the second confidence corresponding to each intended service based on the first confidence and the intention value obtained by the transceiver unit 101 when there is no intended service with a first confidence greater than or equal to the first confidence threshold; the processing unit 102 is also used to determine that the intended service with a second confidence greater than or equal to the second confidence threshold is the user intention.
[0153] In some feasible examples, the processing unit 102 is further configured to determine that the default intended service is the user intention when it is determined that there is no intended service with a second confidence level greater than or equal to a second confidence threshold.
[0154] In some feasible examples, the processing unit 102 is specifically used to input the voice text corresponding to the voice data received by the transceiver unit 101 and the intention value obtained by the transceiver unit 101 into a pre-configured intention recognition model to determine the first confidence corresponding to each intention service.
[0155] In some feasible examples, the transceiver unit 101 is further used to receive an intent setting request sent by the first device; the processing unit 102 is further used to respond to the intent setting request received by the transceiver unit 101, and when it is determined that the personalized intent list corresponding to the first account is not saved, control the transceiver unit 101 to send a default intent list to the first device; the transceiver unit 101 is further used to receive a personalized intent list sent by the first device; the processing unit 102 is further used to establish a correspondence between the first account and the personalized intent list received by the transceiver unit 101, and save the personalized intent list.
[0156] Among them, all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module, and its role will not be repeated here.
[0157] Of course, the server 400 provided in the embodiment of the present application includes but is not limited to the above modules. For example, the server 400 may further include a storage unit 103. The storage unit 103 may be used to store the program code of the write server 400, and may also be used to store data generated by the write server 400 during operation, such as data in a write request.
[0158] As an example, combining Figure 3 The functions implemented by the communication control module 201 and the semantic control module 202 in the server 400 are similar to those implemented by the communication control module 201 and the semantic control module 202 in the server 400. Figure 8 The functions of the transceiver unit 101 are the same as those of the intention recognition module 203 and the business system module 204. Figure 8 The function of the processing unit 102 is the same as that of the data storage module 205. Figure 8 The function of the storage unit 103 in is the same.
[0159] This embodiment of the present application further provides a server, which may include a memory and one or more processors. The memory and processors are coupled. The memory is configured to store computer program code, which includes computer instructions. When the processors execute the computer instructions, the server may perform the functions or steps performed by server 400 in the above-described method embodiment.
[0160] The present application also provides a chip system, which can be applied to the server 400 in the above embodiment. Figure 9 As shown, the chip system includes at least one processor 1501 and at least one interface circuit 1502. The processor 1501 can be the processor in the above-mentioned server 400. The processor 1501 and the interface circuit 1502 can be interconnected via a line. The processor 1501 can receive and execute computer instructions from the memory of the above-mentioned server 400 through the interface circuit 1502. When the computer instructions are executed by the processor 1501, the server 400 can execute the various steps performed by the server 400 in the above-mentioned embodiment. Of course, the chip system can also include other discrete components, which are not specifically limited in this embodiment of the present application.
[0161] The embodiment of the present application also provides a computer-readable storage medium for storing computer instructions executed by the above-mentioned server 400.
[0162] The embodiment of the present application also provides a computer program product, including computer instructions executed by the above-mentioned server 400.
[0163] like Figure 10 As shown, an embodiment of the present application provides a schematic structural diagram of a television set 1. The television set 1 includes a transceiver unit 1001 and a processing unit 1002.
[0164] The transceiver unit 1001 is used to send voice data to the server; wherein, the account currently logged in by the first device is the first account, and the first account corresponds to a personalized intent list, and the personalized intent list includes intent values corresponding to different intent services set by the user; the transceiver unit 1001 is also used to receive business data sent by the server; wherein, the business data includes business data containing voice text in the user's intent, and the user's intention is determined based on the voice text and intention value corresponding to the voice data, and the user's intention is any one of the different intent services.
[0165] In some feasible examples, the identification device also includes a processing unit 1002; the transceiver unit 1001 is also used to send an intent setting request to the server; the transceiver unit 1001 is also used to receive a default intent list sent by the server; the processing unit 1002 is also used to control the transceiver unit 1001 to send a personalized intent list to the server in response to the user's setting operation on the default intent list received by the transceiver unit 1001.
[0166] Among them, all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module, and its role will not be repeated here.
[0167] Of course, the television set 1 provided in the embodiment of the present application includes but is not limited to the above modules. For example, the television set 1 may further include a storage unit 1003. The storage unit 1003 may be used to store the program code for writing the television set 1, and may also be used to store data generated during the operation of the television set 1, such as data in a write request.
[0168] The present application also provides an electronic device, which may include a memory and one or more processors. The memory and processor are coupled. The memory is used to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device may perform the functions or steps performed by the electronic device (e.g., television 1) in the above method embodiment.
[0169] The present application also provides a chip system, which can be applied to the television 1 in the above embodiment. Figure 11 As shown, the chip system includes at least one processor 1601 and at least one interface circuit 1602. The processor 1601 can be the processor in the above-mentioned television 1. The processor 1601 and the interface circuit 1602 can be interconnected via a line. The processor 1601 can receive and execute computer instructions from the memory of the above-mentioned television 1 through the interface circuit 1602. When the computer instructions are executed by the processor 1601, the television 1 can execute the various steps performed by the television 1 in the above-mentioned embodiment. Of course, the chip system can also include other discrete components, which are not specifically limited in this embodiment of the present application.
[0170] The embodiment of the present application further provides a computer-readable storage medium for storing computer instructions executed by the television 1 .
[0171] The embodiment of the present application also provides a computer program product, including computer instructions executed by the above-mentioned television 1.
[0172] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0173] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0174] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0175] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0176] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0177] For ease of explanation, the above description has been made in conjunction with specific embodiments. However, the above discussion of some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are intended to better explain the principles and practical applications, so that those skilled in the art can better use the embodiments and various different variations of the embodiments suitable for specific use considerations.
Claims
1. A method for identifying user intention, characterized in that: include: Receiving voice data sent by a first device; wherein the account currently logged in to the first device is the first account; Obtaining a personalized intent list corresponding to the first account; wherein the personalized intent list includes different intent services set by the user and corresponding intent values; Determine, based on the voice text corresponding to the voice data and the intention value, a first confidence level corresponding to each intended service, and determine that the intended service with the first confidence level greater than or equal to the first confidence threshold is the user intention; if it is determined that there is no intended service with the first confidence level greater than or equal to the first confidence threshold, determine, based on the first confidence level and the intention value, a second confidence level corresponding to each intended service, and determine that the intended service with the second confidence level greater than or equal to the second confidence threshold is the user intention; wherein the intention value is used to reflect the user's preference for the intended service; Calling the service data containing the voice text in the user intention; Send the service data to the first device.
2. The method for identifying user intention according to claim 1, wherein: The identification method further includes: When it is determined that there is no intended service with a second confidence level greater than or equal to a second confidence level threshold, the default intended service is determined to be the user intention.
3. The method for identifying user intention according to claim 1, wherein: The determining, based on the voice text corresponding to the voice data and the intention value, a first confidence level corresponding to each intended service includes: The speech text corresponding to the speech data and the intention value are input into a pre-configured intention recognition model to determine a first confidence level corresponding to each intention service.
4. The method for identifying user intention according to claim 1, wherein: Before receiving the voice data sent by the first device, the recognition method further includes: receiving an intent setting request sent by the first device; In response to the intent setting request, if it is determined that the personalized intent list corresponding to the first account is not saved, sending a default intent list to the first device; Receiving a personalized intent list sent by the first device; Establish a correspondence between the first account and the personalized intent list, and save the personalized intent list.
5. A method for identifying user intention, applied to a first device, characterized in that: include: Sending voice data to the server; wherein the account currently logged in by the first device is the first account, the first account corresponds to a personalized intent list, the personalized intent list includes intent values corresponding to different intent services set by the user; the intent value is used to reflect the user's preference for the intent service; Receive business data sent by the server; wherein, the business data includes business data containing voice text in the user intention, and the user intention is an intention business with a first confidence greater than or equal to a first confidence threshold or, in the absence of an intention business with a first confidence greater than or equal to the first confidence threshold, an intention business with a second confidence greater than or equal to a second confidence threshold; the first confidence corresponding to each intention business is determined according to the voice text corresponding to the voice data and the intention value, and the second confidence corresponding to each intention business is determined according to the first confidence and the intention value.
6. The method for identifying user intention according to claim 5, characterized in that: The identification method further includes: Sending an intent setting request to the server; receiving a default intent list sent by the server; In response to a user setting operation on the default intent list, a personalized intent list is sent to the server.
7. A speech recognition system, characterized in that: The invention comprises a server and an electronic device, wherein the server executes the method for identifying user intention as described in any one of claims 1 to 4 above, and the electronic device executes the method for identifying user intention as described in claim 5 or 6 above.
Citation Information
Patent Citations
Device and method for understanding user intent
CN106663424A