Voice interaction method and electronic equipment

By clarifying ambiguous voice data intentions through user prompts, the method improves the accuracy and relevance of electronic device responses, enhancing voice interaction experience.

CN120319232APending Publication Date: 2025-07-15HONOR DEVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410035628.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing electronic devices cannot accurately analyze the user's true intentions in the voice interaction function, resulting in a large difference between the response results and the user's intentions and a poor user experience.

Method used

By receiving voice data and determining that there is a fuzzy intention in the intention, the confirmation voice data is output to clarify the user's true intention, the reply voice data is used to further confirm the intention, and the response result matching the user's intention is output.

Benefits of technology

It improves the user experience of voice interaction functions, and by clarifying user intentions and outputting more accurate response results, improving the user's real intention confirmation efficiency and response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120319232A_ABST
    Figure CN120319232A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of artificial intelligence, and provides a voice interaction method and electronic equipment. The voice interaction method comprises the following steps: receiving first voice data, and processing the first voice data to obtain an intention corresponding to the first voice data and information of the intention; outputting second voice data under the condition of determining that a fuzzy intention exists in the intention according to the information of the intention; the second voice data is used for confirming a first actual intention of the user for the fuzzy intention; the first actual intention comprises a confirmation intention or a negative intention; receiving reply voice data for the second voice data, and processing the reply voice data to obtain the first actual intention; outputting a response result for the first voice data according to the first actual intention; therefore, the use experience of the voice interaction function can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a voice interaction method and an electronic device. Background Art

[0002] Currently, electronic devices such as mobile phones can support voice interaction functions. The voice interaction function enables users to interact with the electronic device through voice, thereby realizing operations such as voice input, voice control of the electronic device, voice query of information, or voice chat. Exemplarily, the user can control the electronic device to turn on the Bluetooth by inputting the voice command "turn on Bluetooth", or can query weather information by inputting the voice command "what's the weather like in City A today", etc.

[0003] In the process of implementing the voice interaction function, the electronic device needs to first parse the user's intention from the voice data input by the user, and then execute an operation that matches the user's intention. However, in actual applications, the electronic device usually cannot accurately parse the user's true intention from the voice data, and may thus output a response result that is far from the user's true intention, resulting in a poor user experience of the voice interaction function. Summary of the Invention

[0004] Embodiments of this application provide a voice interaction method and an electronic device, which can improve the user experience of the voice interaction function.

[0005] In a first aspect, embodiments of this application provide a voice interaction method, including: receiving first voice data, and processing the first voice data to obtain the intention corresponding to the first voice data and information of the intention; in the case that it is determined according to the information of the intention that there is a fuzzy intention in the intention, outputting second voice data; the second voice data is used to confirm the first actual intention of the user for the fuzzy intention; the first actual intention includes a confirmation intention or a denial intention; receiving reply voice data for the second voice data, and processing the reply voice data to obtain the first actual intention; according to the first actual intention, outputting a response result for the first voice data.

[0006] Wherein, the intention corresponding to the first voice data may refer to the user demand expressed by the first voice data.

[0007] Exemplarily, the information of the intention may include the number of intentions, confidence, and slot information.

[0008] The number of intentions can be used to represent the number of user demands expressed by the first voice data.

[0009] The confidence of the intention can be used to represent the probability that the user demand expressed by the first voice data belongs to this intention.

[0010] A slot can refer to a keyword field related to an intent. Exemplarily, slots can include different types such as time slots, location slots, or entity slots. An entity slot can refer to a slot that takes an application as the value object. It can be understood that one intent can correspond to multiple slots. The slots corresponding to different intents can be the same or different.

[0011] Slot information can be used to represent the value of a slot, and the slot information can help an electronic device better understand the user's intent.

[0012] In any of the following cases, it can be determined that there is a fuzzy intent in the intent corresponding to the first voice data:

[0013] Case 1, the intent corresponding to the first voice data is a single intent, and the confidence of this single intent is within the first confidence interval.

[0014] Case 2, the first voice data corresponds to multiple intents.

[0015] Optionally, in the case where the first actual intent is a confirmation intent, it means that the user has confirmed the fuzzy intent, that is, it means that the fuzzy intent is the user's true intent.

[0016] Optionally, in the case where the first actual intent is a denial intent, it means that the user has denied the fuzzy intent, that is, it means that the fuzzy intent is not the user's true intent.

[0017] According to the voice interaction method provided by the embodiments of the present application, when there is a fuzzy intent in the received first voice data, a second voice data for confirming the user's first actual intent for the fuzzy intent is output, and the true intent of the user is further clarified according to the reply voice data for the second voice data, so that a response result matching the true intent of the user can be output based on the first actual intent, and the use experience of the voice interaction function is improved.

[0018] In an optional implementation manner of the first aspect, when it is determined that there is a fuzzy intent in the intent according to the information of the intent, outputting the second voice data includes: when the intent is a single intent and the confidence of the single intent is within the first confidence interval, determining that there is a fuzzy intent in the intent, and outputting the second voice data based on the single intent and the information of the single intent.

[0019] According to the voice interaction method provided by the embodiments of the present application, when the intent corresponding to the first voice data is a single intent and the confidence of the single intent is within the first confidence interval, a second voice data for confirming the user's first actual intent for this single intent can be output to further clarify the true intent of the user, so that a response result matching the true intent of the user can be output based on the first actual intent, and the use experience of the voice interaction function is improved.

[0020] In an alternative implementation of the first aspect, when it is determined that there is an ambiguous intention in the intention based on the information of the intention, the second voice data is output, including: when the intention is multiple intentions, it is determined that there is an ambiguous intention in the intention, and the second voice data is output based on one target intention with the highest confidence level among the multiple intentions and the information of the target intention.

[0021] According to the voice interaction method provided by the embodiments of the present application, when the intention corresponding to the first voice data is multiple intentions, the second voice data for confirming the user's first actual intention for the target intention among the multiple intentions can be output to further clarify the user's true intention, so that a response result matching the user's true intention can be output based on the first actual intention, improving the usage experience of the voice interaction function.

[0022] In an alternative implementation of the first aspect, according to the first actual intention, the response result for the first voice data is output, including: when the first actual intention is a confirmation intention, according to the slot information of the ambiguous intention, it is determined whether there is an entity slot corresponding to multiple application programs in the slots corresponding to the ambiguous intention; when there is no entity slot, according to the ambiguous intention and the slot information of the ambiguous intention, the target operation matching the ambiguous intention is executed, and the execution result of the target operation is output.

[0023] In an alternative implementation of the first aspect, after determining whether there is an entity slot corresponding to multiple application programs in the slots corresponding to the ambiguous intention according to the slot information of the ambiguous intention, it further includes: when there is an entity slot, determining a target entity from the multiple application programs corresponding to the entity slot; outputting the third voice data; the third voice data is used to confirm the user's second actual intention for the target entity; the second actual intention includes a confirmation intention or a denial intention; receiving the reply voice data for the third voice data, and processing the reply voice data for the third voice data to obtain the second actual intention; outputting the response result according to the ambiguous intention and the second actual intention.

[0024] Optionally, when the second actual intention is a confirmation intention, it means that the user has confirmed the target entity, that is, it means that the target entity conforms to the user's true intention.

[0025] Optionally, when the second actual intention is a denial intention, it means that the user has denied the target entity, that is, it means that the target entity does not conform to the user's true intention.

[0026] In an alternative implementation of the first aspect, outputting a response result according to the fuzzy intention and the second actual intention includes: when the second actual intention is a confirmation intention, performing a target operation matching the fuzzy intention according to the fuzzy intention and the second actual intention, and outputting an execution result of the target operation.

[0027] In an alternative implementation of the first aspect, outputting a response result according to the fuzzy intention and the second actual intention includes: when the second actual intention is a denial intention, outputting voice prompt data; the voice prompt data is used to prompt not to perform a target operation matching the fuzzy intention.

[0028] According to the voice interaction method provided by the embodiments of the present application, when there is an entity slot corresponding to multiple application programs in the slot corresponding to the fuzzy intention, by reconfirming the user's second actual intention for the target entity in the entity slot, the user's true intention can be further refined, so that a response result more matching the user's true intention can be output based on the fuzzy intention and the second actual intention, further improving the usage experience of the voice interaction function.

[0029] In an alternative implementation of the first aspect, determining a target entity from multiple application programs corresponding to the entity slot includes: determining a target entity from multiple application programs corresponding to the entity slot according to the installation status and the last startup time of the multiple application programs corresponding to the entity slot.

[0030] In an alternative implementation of the first aspect, determining a target entity from multiple application programs corresponding to the entity slot according to the installation status and the last startup time of the multiple application programs corresponding to the entity slot includes: determining the application program with the installation status of installed and the latest last startup time among the multiple application programs corresponding to the entity slot as the target entity.

[0031] Wherein, the last startup time of the application program refers to the one of the multiple startup times of the application program that is the closest to the current moment.

[0032] According to the voice interaction method provided by the embodiments of the present application, by determining the application program with the installation status of installed and the latest last startup time among the multiple application programs corresponding to the entity slot as the target entity, since the application program with the latest last startup time is the program that the user used last among the multiple application programs corresponding to the entity slot, therefore, by determining this application program as the target application program, the confirmation efficiency of the user's true intention can be improved, thereby improving the response speed of the voice data and further improving the usage experience of the voice interaction function.

[0033] In an alternative implementation of the first aspect, according to the first actual intention, a response result for the first voice data is output, including: when the first actual intention is a denial intention, voice prompt data is output; the voice prompt data is used to prompt not to perform the target operation matching the ambiguous intention.

[0034] According to the voice interaction method provided by the embodiments of the present application, by outputting, when it is confirmed that the ambiguous intention is not the user's true intention, the voice prompt data for prompting not to perform the target operation matching the ambiguous intention, the user can learn the response status of the terminal device for the first voice data, thereby further improving the usage experience of the voice interaction function.

[0035] In an alternative implementation of the first aspect, after obtaining the intention corresponding to the first voice data and the information of the intention, it further includes: when it is determined according to the information of the intention that there is no ambiguous intention in the intention and there is an entity slot corresponding to multiple application programs in the slots corresponding to the intention, determining a target entity from the multiple application programs corresponding to the entity slot; outputting fourth voice data; the fourth voice data is used to confirm the third actual intention of the user for the target entity; the third actual intention includes a confirmation intention or a denial intention; receiving the reply voice data for the fourth voice data, and processing the reply voice data for the fourth voice data to obtain the third actual intention; according to the intention corresponding to the first voice data and the third actual intention, outputting a response result for the first voice data.

[0036] Among them, the fourth voice data can be used to confirm the third actual intention of the user for the target entity.

[0037] Exemplarily, the third actual intention includes a confirmation intention or a denial intention.

[0038] Optionally, when the third actual intention is a confirmation intention, it means that the user has confirmed the target entity, that is, it means that the target entity conforms to the user's true intention.

[0039] Optionally, when the third actual intention is a denial intention, it means that the user has denied the target entity, that is, it means that the target entity does not conform to the user's true intention.

[0040] According to the voice interaction method provided by the embodiments of the present application, by outputting the fourth voice data for confirming the third actual intention of the user for the target entity when there is no ambiguous intention in the intention corresponding to the first voice data but there is an entity slot corresponding to multiple application programs in the slots corresponding to the intention, the user's true intention is further refined, so that a response result more matching the user's true intention can be output based on the third intention, improving the usage experience of the voice interaction function.

[0041] Second aspect, an embodiment of the present application provides an electronic device, including: one or more processors; one or more memories; the one or more memories store one or more computer-executable programs, and the one or more computer-executable programs include instructions, when the instructions are executed by the one or more processors, the electronic device is caused to execute each step in the voice interaction method according to any implementation manner of the first aspect as described above.

[0042] Third aspect, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium stores a computer-executable program, and when the computer-executable program is called by an electronic device, the electronic device is caused to execute each step in the voice interaction method according to any implementation manner of the first aspect as described above.

[0043] Fourth aspect, an embodiment of the present application provides a computer-executable program product, when the computer-executable program product runs on an electronic device, the electronic device is caused to execute each step in the voice interaction method according to any implementation manner of the first aspect as described above.

[0044] Fifth aspect, an embodiment of the present application provides a chip system, the chip system is applied to an electronic device, the chip system includes a processor, the processor is coupled to a memory, the memory is used to store computer program instructions, when the processor calls the computer program instructions, the electronic device is caused to implement each step in the voice interaction method according to any implementation manner of the first aspect as described above. The chip system may be a single chip or a chip module composed of multiple chips.

[0045] It can be understood that the beneficial effects of the second aspect to the fifth aspect as described above can refer to the relevant descriptions in the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0047] Figure 2 is a schematic software architecture diagram of an electronic device provided by an embodiment of the present application;

[0048] Figure 3 is a schematic flowchart of a voice interaction method provided by an embodiment of the present application;

[0049] Figure 4 is a schematic diagram of the interaction logic between modules in the system architecture of an electronic device during the implementation process of a voice interaction method provided by an embodiment of the present application;

[0050] Figure 5During the implementation of a voice interaction method provided by another embodiment of the present application, a schematic diagram of the interaction logic between modules in the system architecture of an electronic device;

[0051] Figure 6 A specific implementation flowchart of S34 in a voice interaction method provided by an embodiment of the present application;

[0052] Figure 7 During the implementation of a voice interaction method provided by another embodiment of the present application, a schematic diagram of the interaction logic between modules in the system architecture of an electronic device;

[0053] Figure 8 A schematic flowchart of a voice interaction method provided by another embodiment of the present application;

[0054] Figure 9 During the implementation of a voice interaction method provided by another embodiment of the present application, a schematic diagram of the interaction logic between modules in the system architecture of an electronic device;

[0055] Figure 10 A schematic diagram of the structure of a DM module provided by an embodiment of the present application. Detailed implementation manners

[0056] It should be noted that the terms used in the implementation manner part of the embodiments of the present application are only used to explain the specific embodiments of the present application, rather than intended to limit the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B can mean A or B; herein, "and / or" is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two, and "at least one", "one or more" mean one, two or more than two.

[0057] Hereinafter, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0058] References to "one embodiment" or "some embodiments" etc. described in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0059] Currently, electronic devices such as mobile phones can support voice interaction functions. The voice interaction function enables users to interact with the electronic device through voice, thereby realizing operations such as voice input, voice control of the electronic device, voice query of information, or voice chat. Exemplarily, the user can control the electronic device to turn on Bluetooth by inputting the voice command "turn on Bluetooth", or can realize the query of weather information by inputting the voice command "what's the weather like in City A today", etc.

[0060] In the process of implementing the voice interaction function, the electronic device needs to first parse the user's intention from the voice data input by the user, and then can execute the operation that matches the user's intention. For example, the electronic device needs to first parse the user's intention from the voice data "what's the weather like in City A today" as "query the weather in City A today", and then execute the weather query operation. However, in actual applications, the electronic device usually cannot accurately parse the user's true intention from the voice data, and thus may output a response result that is very different from the user's true intention, resulting in a poor user experience of the voice interaction function.

[0061] Exemplarily, in some scenarios, when the user inputs a piece of voice data "what's the weather like in City A today, want to take a photo", the electronic device may parse two intentions, "query weather" and "take a photo", from this piece of voice data, but cannot determine whether the user's true intention is to query the weather or take a photo. In such a voice interaction scenario where the voice data contains multiple intentions, the electronic device will output a response result indicating that it cannot understand the user's intention, thereby resulting in a poor user experience of the voice interaction function.

[0062] In some other scenarios, when a user inputs a voice data "Is it cold when climbing the mountain tomorrow", the electronic device may not be able to accurately parse out that the user's intention is to "query the weather", that is, the confidence level of the electronic device parsing out the intention of "querying the weather" from this voice data is relatively low. In such a voice interaction scenario where the voice data only includes a single intention but the confidence level of the single intention is relatively low, the electronic device will also output a response result indicating that it does not understand the user's intention, thus resulting in a poor user experience of the voice interaction function.

[0063] In still some other scenarios, when a user inputs a voice data "Open Ghost Blows Out the Light", since "Ghost Blows Out the Light" may be a movie or a novel, that is, "Ghost Blows Out the Light" may involve multiple application entities, such as a video application and a reading application, the electronic device cannot accurately determine whether the user's true intention is to open the "Ghost Blows Out the Light" movie in the video application or open the "Ghost Blows Out the Light" novel in the reading application. In such a voice interaction scenario where the voice data involves multiple application entities, the electronic device will also output a response result indicating that it does not understand the user's intention, thus resulting in a poor user experience of the voice interaction function.

[0064] In view of this, embodiments of the present application provide a voice interaction method and an electronic device. When there is a fuzzy intention in the received first voice data, a second voice data is output for confirming the first actual intention of the user for the fuzzy intention, and the true intention of the user is further clarified according to the reply voice data for the second voice data, so that a response result matching the true intention of the user can be output based on the first actual intention, improving the user experience of the voice interaction function.

[0065] The voice interaction method provided by embodiments of the present application can be applied to electronic devices such as mobile phones, tablet computers, wearable devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). Embodiments of the present application do not limit the specific type of the electronic device.

[0066] Exemplarily, please refer to Figure 1 , which is a schematic structural diagram of an electronic device provided by embodiments of the present application.

[0067] As Figure 1As shown in the figure, the electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0068] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0069] Exemplarily, the processor 110 may be used to execute the voice interaction method in the embodiments of the present application.

[0070] The controller may be the nerve center and command center of the electronic device. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.

[0071] A memory can also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can hold the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0072] The electronic device can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc. For example, collecting voice data or outputting the response results of voice data, etc.

[0073] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals.

[0074] In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.

[0075] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device can output the voice interaction result in the voice interaction logic through the speaker 170A.

[0076] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device answers a call or a voice message, the voice can be received by bringing the receiver 170B close to the human ear.

[0077] The microphone 170C, also known as the "microphone", "transmitter", is used to convert a sound signal into an electrical signal. When making a call, sending a voice message, or inputting voice data, the user can speak close to the microphone 170C with the mouth to input the sound signal into the microphone 170C. The electronic device can be provided with at least one microphone 170C.

[0078] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0079] It can be understood that the above is an exemplary description of the structure of the electronic device. It should be understood that in other embodiments, the electronic device may include more or fewer components than those shown in the figure, or certain components may be combined, or certain components may be split, or different component arrangements may be adopted. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0080] The software system of the electronic device may adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In this embodiment of the application, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software architecture of the electronic device.

[0081] Please refer to Figure 2 , which is a schematic diagram of the software architecture of an electronic device provided in this embodiment of the application. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. For example, the Android system can be divided into four layers, from top to bottom, namely the application layer, the application framework layer, the Android runtime and the system libraries, and the kernel layer.

[0082] The application layer may include a series of application packages. For example, it may include a voice assistant, a voice processing application, etc.

[0083] The voice assistant and the voice processing application can cooperate with each other to jointly implement the voice interaction method provided in this embodiment of the application.

[0084] Among them, the voice assistant can be used to be responsible for the interaction part in the voice interaction logic. For example, the acquisition of voice data input by the user, the output of the response result corresponding to the voice data (including text display or audio output), etc.

[0085] The voice processing application can be used to be responsible for the part of voice data processing in the voice interaction logic. For example, the parsing of the user intention corresponding to the voice data, the decision-making of the target operation or response result matching the user's intention, etc.

[0086] Exemplarily, the voice processing application may include a voice activity detection (VAD) module, a voice distribution module, an automatic speech recognition (ASR) module, a natural language understanding (NLU) module, and a dialog management (DM) module.

[0087] Among them, the VAD module can be used to identify the active part in the voice data input by the user, that is, to determine whether there is a voice part in an audio segment, exclude the non-voice part in the audio (such as silence or noise, etc.), and transmit the voice part to the voice distribution module. That is, the VAD module can be used to detect when there is a voice signal and when there is no voice signal.

[0088] The voice distribution module can be used to implement data interaction among the VAD module, ASR module, NLU module, and DM module. That is, the voice distribution module can be equivalent to a data scheduling center in the voice interaction logic.

[0089] The ASR module can be used to perform speech recognition on the voice data input by the user, so as to convert the voice data input by the user into the corresponding text.

[0090] The NLU module can be used to perform semantic analysis on the text corresponding to the voice data input by the user, and obtain the intent and information of the intent corresponding to the voice data input by the user, such as the confidence level of the intent or slot information, etc.

[0091] The DM module can be used to manage and organize the dialogue process in the voice interaction logic. For example, it can be used to decide whether to respond to the voice data input by the user, or how to respond to the voice data input by the user, etc., based on the intent and information of the intent corresponding to the voice data obtained by the NLU module.

[0092] It should be noted that the interaction logic among the various modules in the voice processing application during the implementation process of the voice interaction method will be introduced in detail in the subsequent embodiments and will not be elaborated here for the time being.

[0093] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.

[0094] Android Runtime includes core libraries and virtual machines. Android runtime is responsible for the scheduling and management of the Android system.

[0095] The core libraries contain two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core libraries of Android.

[0096] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.

[0097] The system library may include multiple functional modules. For example: surface manager, media library, 3D graphics processing library, 2D graphics engine, etc.

[0098] The kernel layer is the layer between hardware and software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.

[0099] The voice interaction method provided by the embodiments of the present application can be applied to the above-mentioned electronic device. The voice interaction method provided by the embodiments of the present application will be described in detail below in combination with the above structure of the electronic device.

[0100] Please refer to Figure 3 , which is a schematic flowchart of a voice interaction method provided by the embodiments of the present application. In some embodiments of the present application, the voice interaction method may include S31 to S34, which are described in detail as follows:

[0101] S31, receive the first voice data, and process the first voice data to obtain the intent corresponding to the first voice data and the information of the intent.

[0102] In the embodiments of the present application, when the voice assistant is awakened, the voice assistant can control the microphone to collect voice data, and can transmit the voice data collected by the microphone to the voice processing application. Exemplarily, the voice assistant can transmit the first voice data collected by the microphone to the voice processing application. After receiving the first voice data, the voice processing application can process the first voice data to determine the intent corresponding to the first voice data and the information of the intent.

[0103] Among them, the intent corresponding to the first voice data may refer to the user's needs expressed by the first voice data.

[0104] Exemplarily, assuming that the first voice data is "What's the weather like in City A today", it indicates that the user has a need to query the weather. Therefore, the voice processing application can determine that the intent corresponding to the first voice data is "query weather".

[0105] Exemplarily, the information of the intent may include the number of intents, confidence, and slot information.

[0106] Among them, the number of intents can be used to represent the number of user needs expressed by the first voice data.

[0107] The confidence of the intent can be used to represent the probability that the user needs expressed by the first voice data belong to this intent.

[0108] A slot can refer to a keyword field related to an intent. Exemplarily, slots can include different types such as time slots, location slots, or entity slots. Exemplarily, an entity slot can refer to a slot with an application as the value-taking object. It can be understood that one intent can correspond to multiple slots. Slots corresponding to different intents can be the same or different.

[0109] Slot information can be used to represent the values of slots, and the slot information can help the electronic device better understand the user's intent. Exemplarily, assuming that the slots of the intent of "querying weather" parsed from the first voice data "What's the weather like in City A today" include a location slot and a time slot, the value of the location slot is "City A", and the value of the time slot is "today", then the slot information of the intent of "querying weather" can include: Location: City A, Time: today.

[0110] Exemplarily, the following combines Figure 4 and Figure 5 , and illustrates the interaction logic between modules in the system architecture of the electronic device during the implementation process of the voice interaction method provided in the embodiments of the present application.

[0111] Among them, Figure 4 It can correspond to the implementation logic of the voice interaction method in a single-intent scenario where the confidence level is within the first confidence interval (for example, [0.5, 0.8]). Figure 5 It can correspond to the implementation logic of the voice interaction method in a multi-intent scenario.

[0112] Exemplarily, in combination with Figure 4 in (a) or Figure 5 in (a), when specifically implemented, the voice processing application can receive the first voice data from the voice assistant through the VAD module; after receiving the first voice data, the VAD module can transmit the first voice data to the voice distribution module. The voice distribution module can transmit the first voice data to the ASR module. The ASR module can perform speech recognition on the first voice data to obtain the first text corresponding to the first voice data, and return the first text to the voice distribution module. The voice distribution module can transmit the first text corresponding to the first voice data to the NLU module. The NLU module can perform semantic recognition on the first text to parse out the intent and the information of the intent corresponding to the first voice data from the first text, and can return the intent and the information of the intent corresponding to the first voice data to the voice distribution module. The voice distribution module can transmit the intent and the intent information corresponding to the first voice data to the DM module. The DM module can determine whether there is a fuzzy intent in the intent corresponding to the first voice data according to the information of the intent corresponding to the first voice data.

[0113] Optionally, in any of the following cases, the DM module may determine that there is an ambiguous intention in the intention corresponding to the first voice data:

[0114] Case 1: The intention corresponding to the first voice data is a single intention, and the confidence of this single intention is within the first confidence interval.

[0115] Among them, the first confidence interval may be a pre-configured interval for limiting the confidence of ambiguous intentions. Exemplarily, the first confidence interval may be [0.5, 0.8].

[0116] In this case, since the confidence of the single intention corresponding to the first data is within the first confidence interval, the DM module may consider that the user's intention is not clear. Based on this, the DM module may determine this single intention as an ambiguous intention.

[0117] Case 2: The first voice data corresponds to multiple intentions.

[0118] In this case, since the first voice data corresponds to multiple intentions, the DM module may consider that the user's intention is not clear regardless of whether the confidence of these multiple intentions is within the first confidence interval. Based on this, the DM module may determine these multiple intentions as ambiguous intentions.

[0119] Optionally, when the intention corresponding to the first voice data is single and the confidence of this single intention is greater than the upper limit value of the first confidence interval (e.g., 0.8), the DM module may consider that the user's intention is clear and determine that there is no ambiguous intention in the intention corresponding to the first voice data.

[0120] In an optional implementation manner, when there is an ambiguous intention in the intention corresponding to the first voice data, the electronic device may execute S32 to S34.

[0121] In another alternative implementation, when there is no ambiguous intention in the intention corresponding to the first voice data, the electronic device may perform a target operation that matches the intention corresponding to the first voice data based on the intention corresponding to the first voice data and the information of the intention, and output the execution result of the target operation. Exemplarily, assuming that the confidence level of the intention of "querying the weather" corresponding to the first voice data "query the weather in City A today" is 1, and the slot information includes: location: City A, time: today; then the DM module may determine that there is no ambiguous intention in the intention corresponding to the first voice data, and may call the weather application to perform a weather query operation. Assuming that the weather query result returned by the weather application is "It is sunny in City A today", the DM module may send a first response instruction to the voice distribution module, and the first response instruction may be used to instruct the voice assistant to output the weather query result of "It is sunny in City A today". The voice distribution module may transmit the first response instruction to the voice assistant. The voice assistant may broadcast the voice data of "It is sunny in City A today" through the speaker, and may also display the text of "It is sunny in City A today" on the screen.

[0122] S32. When it is determined that there is an ambiguous intention in the intention according to the information of the intention, output the second voice data.

[0123] Wherein, the second voice data may be used to confirm the user's first actual intention for the ambiguous intention.

[0124] Exemplarily, the first actual intention may include a confirmation intention or a denial intention.

[0125] Optionally, when the first actual intention is a confirmation intention, it means that the user has confirmed the ambiguous intention, that is, it means that the ambiguous intention is the user's true intention.

[0126] Optionally, when the first actual intention is a denial intention, it means that the user has denied the ambiguous intention, that is, it means that the ambiguous intention is not the user's true intention.

[0127] In a specific implementation, in the above case 1, that is, when the intention corresponding to the first voice data is a single intention and the confidence level of the single intention is within the first confidence interval, S32 may specifically include:

[0128] Output the second voice data based on the single intention and the slot information of the single intention.

[0129] Specifically, the DM module can generate a first follow-up instruction based on the single intent and the slot information of the single intent, and return the first follow-up instruction to the voice distribution module. The first follow-up instruction may carry first follow-up information for confirming whether the single intent is the true intent of the user, and the first follow-up instruction can be used to instruct the voice assistant to output the first follow-up information. The voice distribution module can transmit the first follow-up instruction to the voice assistant. After receiving the first follow-up instruction, the voice assistant can generate second voice data corresponding to the first follow-up information and broadcast the second voice data through the speaker. The first follow-up information can also be displayed on the screen.

[0130] Exemplarily, in combination with Figure 4 in (a), assume that the first voice data input by the user is "Is it cold when climbing the mountain tomorrow", the NLU module parses an intent of "query weather" from the first voice data, and determines that the confidence of the intent of "query weather" is 0.5, and the slot information is: location: XX (such as the current location), time: tomorrow; then the DM module can generate the first follow-up information "Confirm to query the weather tomorrow" for confirming whether "query weather" is the true intent of the user based on the intent of "query weather" and the slot information of this intent, and send a first follow-up instruction carrying the first follow-up information "Confirm to query the weather tomorrow" to the voice distribution module. The voice distribution module can transmit the first follow-up instruction carrying the first follow-up information "Confirm to query the weather tomorrow" to the voice assistant. The voice assistant can generate second voice data with the content of "Confirm to query the weather tomorrow" according to the first follow-up information and broadcast the second voice data "Confirm to query the weather tomorrow" through the speaker. The voice assistant can also display the first follow-up information "Confirm to query the weather tomorrow" on the screen.

[0131] In another specific implementation manner, in the above case 2, that is, when the first voice data corresponds to multiple intents, S32 may specifically include:

[0132] Output the second voice data based on the target intent with the highest confidence among the multiple intents and the slot information of the target intent.

[0133] Optionally, when there are multiple intents with the highest confidence among the intents corresponding to the first voice data, the DM module can determine any one of the multiple intents with the highest confidence as the target intent.

[0134] Specifically, the DM module can generate a first follow-up instruction based on the target intention with the highest confidence among multiple intentions corresponding to the first voice data and the slot information of the target intention, and return the first follow-up instruction to the voice distribution module. The voice distribution module can transmit the first follow-up instruction to the voice assistant. Among them, the first follow-up instruction can carry first follow-up information for confirming whether the target intention is the true intention of the user, and the first follow-up instruction can be used to instruct the voice assistant to output the first follow-up information. After receiving the first follow-up instruction, the voice assistant can generate second voice data corresponding to the first follow-up information and broadcast the second voice data through the speaker. The first follow-up information can also be displayed on the screen.

[0135] Exemplarily, in combination with Figure 5 in (a), assume that the first voice data input by the user is "What's the weather like in City A today and I want to take pictures". The NLU module parses two intentions, "query weather" and "take pictures", from the first voice data, and determines that the confidence of the "query weather" intention is 0.7, and the slot information is: location: City A, time: today; the confidence of the "take pictures" intention is 0.6. Since the confidence of the "query weather" intention is the highest, the DM module can generate the first follow-up information "Do you confirm to query the weather in City A today" for confirming whether "query the weather in City A today" is the true intention of the user based on the "query weather" intention and the slot information of this intention, and send a first follow-up instruction carrying the first follow-up information "Do you confirm to query the weather in City A today" to the voice distribution module. The voice distribution module can transmit the first follow-up instruction carrying the first follow-up information "Do you confirm to query the weather in City A today" to the voice assistant. The voice assistant can generate second voice data with the content of "Do you confirm to query the weather in City A today" according to the first follow-up information and broadcast the second voice data "Do you confirm to query the weather in City A today" through the speaker. In addition, the voice assistant can also display the first follow-up information "Do you confirm to query the weather tomorrow" on the screen.

[0136] S33, receive the reply voice data for the second voice data, and process the reply voice data to obtain the first actual intention.

[0137] In this embodiment, after outputting the second voice data through the speaker, the voice assistant can control the microphone to collect the reply voice data for the second voice data, and can transmit the reply voice data collected by the microphone to the voice processing application. The voice processing application can process the reply voice data to parse the first actual intention from the reply voice data.

[0138] Exemplarily, the voice assistant can determine the voice data collected within the first preset duration after outputting the second voice data as the reply voice data for the second voice data. The first preset duration can be, for example, 4 seconds.

[0139] Exemplarily, in combination with Figure 4 (b) in Figure 5 (b), when specifically implemented, the voice processing application can receive the reply voice data from the voice assistant through the VAD module; the VAD module can transmit the received reply voice data to the voice distribution module. The voice distribution module can transmit the reply voice data to the ASR module. The ASR module can perform speech recognition on the reply voice data to obtain the reply text corresponding to the reply voice data, and return the reply text to the voice distribution module. The voice distribution module can transmit the reply text to the NLU module. The NLU module can perform semantic recognition on the reply text to parse out the first actual intention corresponding to the reply voice data from the reply text.

[0140] S34. According to the first actual intention, output a response result for the first voice data.

[0141] Exemplarily, in combination with Figure 4 (b) in Figure 5 (b), after the NLU module determines the first actual intention corresponding to the reply voice data, it can return the first actual intention to the voice distribution module. The voice distribution module can transmit the first actual intention corresponding to the reply voice data to the DM module. The DM module can determine whether to execute a target operation matching the fuzzy intention according to the first actual intention corresponding to the reply voice data.

[0142] In an optional implementation manner, when the first actual intention is a confirmation intention, the DM module can determine to execute a target operation matching the fuzzy intention. Based on this, the DM module can execute the target operation matching the fuzzy intention according to the fuzzy intention and the slot information of the fuzzy intention, and output the execution result of the target operation.

[0143] Specifically, the DM module can execute the target operation by calling the target application corresponding to the target operation to obtain the execution result of the target operation, and return the execution result of the target operation to the voice distribution module. The voice distribution module can transmit the execution result of the target operation to the voice assistant. After receiving the execution result of the target operation, the voice assistant can broadcast the voice data corresponding to the execution result of the target operation through the speaker, and can also display the execution result of the target operation on the screen.

[0144] Exemplarily, in combination with Figure 4In (b) thereof, assuming that after the voice assistant outputs the second voice data "Do you confirm to query the weather tomorrow?" through the speaker, the voice assistant collects the reply voice data "Confirm" for the second voice data through the microphone, the voice assistant can transmit the reply voice data "Confirm" to the VAD module in the voice processing application. The VAD module can transmit the reply voice data "Confirm" to the voice distribution module. The voice distribution module can transmit the reply voice data "Confirm" to the ASR module. The ASR module can perform speech recognition on the reply voice data "Confirm" to obtain the reply text "Confirm" corresponding to the reply voice data "Confirm", and return the reply text "Confirm" to the voice distribution module. The voice distribution module can transmit the reply text "Confirm" to the NLU module. The NLU module can perform semantic recognition on the reply text "Confirm" and parse out the first actual intention as the confirmation intention from the reply text "Confirm". Based on this, the DM module can execute the weather query operation by calling the weather application. Assuming that the execution result of the weather query operation (i.e., the weather query result) is "Tomorrow is sunny", the DM module can transmit the weather query result "Tomorrow is sunny" to the voice distribution module. The voice distribution module can transmit the weather query result "Tomorrow is sunny" to the voice assistant. The voice assistant can generate the voice data "Tomorrow is sunny" corresponding to the weather query result and broadcast the voice data "Tomorrow is sunny" corresponding to the weather query result through the speaker. In addition, the voice assistant can also display the weather query result "Tomorrow is sunny" on the screen.

[0145] In another alternative implementation, when the first actual intention is the denial intention, the DM module can determine not to execute the target operation matching the ambiguous intention. In this case, the DM module can generate a prompt instruction and transmit the prompt instruction to the voice distribution module. Among them, the prompt instruction can carry prompt information, and the prompt instruction can be used to instruct the voice assistant to output the prompt information. The prompt information can be used to prompt not to execute the target operation matching the ambiguous intention. The voice distribution module can transmit the prompt instruction to the voice assistant. The voice assistant can generate voice prompt data corresponding to the prompt information based on the prompt instruction and can broadcast the voice prompt data through the speaker. In addition, the voice assistant can also display the prompt information on the screen.

[0146] Exemplarily, in combination with Figure 5In (b) of [description above], assume that after the voice assistant outputs the second voice data "Do you confirm to query the weather in City A today?" through the speaker, the voice assistant collects the response voice data "Cancel" for the second voice data through the microphone. Then, the voice assistant can transmit the response voice data "Cancel" to the VAD module in the voice processing application. The VAD module can transmit the response voice data "Cancel" to the voice distribution module. The voice distribution module can transmit the response voice data "Cancel" to the ASR module. The ASR module can perform speech recognition on the response voice data "Cancel" to obtain the response text "Cancel" corresponding to the response voice data "Cancel", and transmit the response text "Cancel" to the voice distribution module. The voice distribution module can transmit the response text "Cancel" to the NLU module. The NLU module can perform semantic recognition on the response text "Cancel" and parse out the first actual intention as a denial intention from the response text "Cancel". Based on this, the DM module can determine not to perform the weather query operation, generate a prompt instruction carrying the prompt message "Okay", and can return the prompt instruction to the voice distribution module. The voice distribution module can transmit the prompt instruction to the voice assistant. The voice assistant can generate the voice prompt data "Okay" corresponding to the prompt message based on the prompt instruction, and can broadcast the voice prompt data "Okay" through the speaker. In addition, the voice assistant can also display the prompt message "Okay" on the screen.

[0147] As can be seen from the above, in the voice interaction method provided by the embodiments of the present application, when there is a fuzzy intention in the received first voice data, the second voice data for confirming the first actual intention of the user for the fuzzy intention is output, and the true intention of the user is further clarified according to the response voice data for the second voice data, so that a response result matching the true intention of the user can be output based on the first actual intention, improving the use experience of the voice interaction function.

[0148] It can be understood that in some scenarios, the first voice data input by the user may not only have a fuzzy intention, but there may also be entity slots corresponding to multiple application programs (i.e., multiple entities) in the slots corresponding to the fuzzy intention. In this scenario, after clarifying the true intention of the user through the first actual intention, the true intention of the user for the multiple entities corresponding to the entity slots can be further confirmed.

[0149] Based on this, in some other embodiments, S34 may specifically include S341 to S346 as Figure 6 shown below, which are described in detail as follows:

[0150] S341, when the first actual intention is a confirmation intention, determine whether there is an entity slot corresponding to multiple application programs in the slot corresponding to the fuzzy intention according to the slot information of the fuzzy intention.

[0151] Among them, the entity slot corresponding to multiple application programs may refer to an entity slot with multiple values.

[0152] Optionally, when there is no entity slot corresponding to multiple application programs in the slots corresponding to the fuzzy intention, the electronic device may execute S342.

[0153] Optionally, when there is an entity slot corresponding to multiple application programs in the slots corresponding to the fuzzy intention, the electronic device may execute S343 to S346.

[0154] S342. When there is no entity slot corresponding to multiple application programs, according to the fuzzy intention and the slot information of the fuzzy intention, execute the target operation that matches the fuzzy intention, and output the execution result of the target operation.

[0155] It can be understood that when the first actual intention of the user for the fuzzy intention is a confirmation intention, and there is no entity slot corresponding to multiple application programs in the slots corresponding to the fuzzy intention, it means that the DM module has already been very clear about the user's true intention. Therefore, the DM module can directly execute the target operation that matches the fuzzy intention according to the fuzzy intention and the slot information of the fuzzy intention, and output the execution result of the target operation.

[0156] Specifically, the DM module may execute the target operation by calling the target application program corresponding to the target operation, obtain the execution result of the target operation, and return the execution result of the target operation to the voice distribution module. The voice distribution module may transmit the execution result of the target operation to the voice assistant. After receiving the execution result of the target operation, the voice assistant may broadcast the voice data corresponding to the execution result of the target operation through the speaker, and may also display the execution result of the target operation on the screen.

[0157] It should be noted that this situation may correspond to Figure 4 the scenario shown in Figure 4 For the relevant description in the embodiment corresponding to (b) therein, details are not described here.

[0158] S343. When there is an entity slot corresponding to multiple application programs, determine the target entity from the multiple application programs corresponding to the entity slot.

[0159] When there is an entity slot corresponding to multiple application programs in the slots corresponding to the fuzzy intention, the DM module may first determine the target entity to be confirmed with the user from the multiple application programs corresponding to the entity slot.

[0160] In an alternative implementation, the DM module can determine the target entity in the following manner: based on the installation status and the most recent startup time of multiple applications corresponding to the entity slot, determine the target entity from the multiple applications corresponding to the entity slot.

[0161] Among them, the most recent startup time of the application can refer to the one startup time that is closest to the current moment among the multiple startup times of the application.

[0162] In a specific implementation, the DM module can determine, among the multiple applications corresponding to the entity slot, the application with the installed status as installed and the latest most recent startup time as the target entity.

[0163] Exemplarily, assume that the slot information of the intent "Open Ghost Blows Out the Light" corresponding to the first voice data "Open Ghost Blows Out the Light" includes: entities: Video APP, Reading APP. Among them, the installation status of the Video APP and the Reading APP are both installed. The most recent startup time of the Video APP is 12:30 on November 30, 2023, and the most recent startup time of the Reading APP is 08:10 on November 27, 2023. Then the DM module can determine the Video APP as the target entity.

[0164] S344, output the third voice data.

[0165] Among them, the third voice data can be used to confirm the user's second actual intent for the target entity.

[0166] Exemplarily, the second actual intent can include a confirmation intent or a denial intent.

[0167] Optionally, in the case where the second actual intent is a confirmation intent, it means that the user has confirmed the target entity, that is, it means that the target entity conforms to the user's true intent.

[0168] Optionally, in the case where the second actual intent is a denial intent, it means that the user has denied the target entity, that is, it means that the target entity does not conform to the user's true intent.

[0169] Specifically, after the DM module determines the target entity, it can generate a second follow-up instruction based on the target entity and return the second follow-up instruction to the voice distribution module. Among them, the second follow-up instruction can carry second follow-up information for confirming whether the target entity conforms to the user's true intent. The second follow-up instruction can be used to instruct the voice assistant to output the second follow-up information. The voice distribution module can transmit the second follow-up instruction to the voice assistant. After receiving the second follow-up instruction, the voice assistant can generate third voice data corresponding to the second follow-up information and broadcast the third voice data through the speaker. It can also display the second follow-up information on the screen.

[0170] Exemplarily, in combination with Figure 7 (a) and (b) in, assume that the first voice data input by the user is "What time is it now, open Ghost Blows Out the Light", and the NLU module determines that the intents corresponding to the first voice data include "query time" and "open Ghost Blows Out the Light", and the confidence of the intent of "query time" is 0.5; the confidence of the intent of "open Ghost Blows Out the Light" is 0.7, and the slot information includes: entity: video APP, reading APP. After one round of confirmation with the user and confirming that the first actual intent of the user for the intent of "open Ghost Blows Out the Light" is the confirmation intent, as Figure 7 shown in (b) in, assume that the DM module determines "video APP" as the target entity, then the DM module can generate a second follow-up information "Confirm to open the Ghost Blows Out the Light movie" for confirming whether the "video APP" conforms to the user's true intent based on the target entity, and send a second follow-up instruction carrying the second follow-up information "Confirm to open the Ghost Blows Out the Light movie" to the voice distribution module. The voice distribution module can transmit the second follow-up instruction carrying the second follow-up information "Confirm to open the Ghost Blows Out the Light movie" to the voice assistant. The voice assistant can generate third voice data with the content of "Confirm to open the Ghost Blows Out the Light movie" according to the second follow-up information, and broadcast the third voice data "Confirm to open the Ghost Blows Out the Light movie" through the speaker. In addition, the voice assistant can also display the second follow-up information "Confirm to open the Ghost Blows Out the Light movie" on the screen.

[0171] S345, receive the reply voice data for the third voice data, and process the reply voice data for the third voice data to obtain the second actual intent.

[0172] In this embodiment, after outputting the third voice data through the speaker, the voice assistant can control the microphone to collect the reply voice data for the third voice data, and can transmit the reply voice data for the third voice data to the voice processing application. The voice processing application can process the reply voice data for the third voice data to determine the second actual intent corresponding to the reply voice data for the third voice data.

[0173] Exemplarily, the voice assistant can determine the voice data collected within the first preset duration after outputting the third voice data as the reply voice data for the third voice data.

[0174] Exemplarily, in combination with Figure 7In (c) of [description], in a specific implementation, the speech processing application can receive the reply speech data for the third speech data through the VAD module; after receiving the reply speech data for the third speech data, the VAD module can transmit the reply speech data for the third speech data to the speech distribution module. The speech distribution module can transmit the reply speech data for the third speech data to the ASR module. The ASR module can perform speech recognition on the reply speech data for the third speech data to obtain the reply text corresponding to the reply speech data for the third speech data, and return the reply text to the speech distribution module. The speech distribution module can transmit the reply text corresponding to the reply speech data for the third speech data to the NLU module. The NLU module can perform semantic recognition on the reply text to parse out the second actual intention corresponding to the reply speech data for the third speech data from the reply text.

[0175] S346, output a response result according to the fuzzy intention and the second actual intention.

[0176] Exemplarily, in combination with Figure 7 In (c) of [description], after the NLU module determines the second actual intention, it can return the second actual intention to the speech distribution module. The speech distribution module can transmit the second actual intention to the DM module. The DM module can determine whether to execute the target operation matching the fuzzy intention according to the second actual intention.

[0177] In an alternative implementation, when the second actual intention is a confirmation intention, the DM module can determine to execute the target operation matching the fuzzy intention. Based on this, S346 can specifically include: execute the target operation matching the fuzzy intention according to the fuzzy intention and the second actual intention, and output the execution result of the target operation.

[0178] Specifically, the DM module can execute the target operation by calling the target application corresponding to the target operation to obtain the execution result of the target operation, and return the execution result of the target operation to the speech distribution module. The speech distribution module can transmit the execution result of the target operation to the voice assistant. After receiving the execution result of the target operation, the voice assistant can broadcast the voice data corresponding to the execution result of the target operation through the speaker, and can also display the execution result of the target operation on the screen.

[0179] Exemplarily, in combination with Figure 7In (c) above, assuming that after the voice assistant outputs the third voice data "Confirm to open the Ghost Blows Out the Light movie" through the speaker, and the reply voice data "Confirm" for the third voice data is collected through the microphone, the voice assistant can transmit the reply voice data "Confirm" to the VAD module in the voice processing application. The VAD module can transmit the reply voice data "Confirm" to the voice distribution module. The voice distribution module can transmit the reply voice data "Confirm" to the ASR module. The ASR module can perform voice recognition on the reply voice data "Confirm" to obtain the reply text "Confirm" corresponding to the reply voice data "Confirm", and return the reply text "Confirm" to the voice distribution module. The voice distribution module can transmit the reply text "Confirm" to the NLU module. The NLU module can perform semantic recognition on the reply text "Confirm" and parse out the second actual intention as the confirmation intention from the reply text "Confirm".

[0180] Based on this, the DM module can perform the operation of opening the Ghost Blows Out the Light movie by calling the video APP. Assuming that the execution result of the operation of opening the Ghost Blows Out the Light movie is "The Ghost Blows Out the Light movie has been opened", the DM module can transmit the execution result "The Ghost Blows Out the Light movie has been opened" to the voice distribution module. The voice distribution module can transmit the execution result "The Ghost Blows Out the Light movie has been opened" to the voice assistant. The voice assistant can generate the voice data "The Ghost Blows Out the Light movie has been opened" corresponding to the execution result, and broadcast the voice data "The Ghost Blows Out the Light movie has been opened" corresponding to the execution result through the speaker. In addition, the voice assistant can also display the execution result "The Ghost Blows Out the Light movie has been opened" on the screen.

[0181] In another alternative implementation, when the second actual intention is the denial intention, the DM module can determine not to perform the target operation matching the fuzzy intention. Based on this, S346 can specifically include: outputting voice prompt data. The voice prompt data can be used to prompt not to perform the target operation matching the fuzzy intention.

[0182] As can be seen from the above, in the voice interaction method provided in this embodiment, when there are entity slots corresponding to multiple application programs in the slot corresponding to the fuzzy intention, by reconfirming the user's second actual intention for the target entity in the entity slot, the user's true intention can be further improved, so that a response result more matching the user's true intention can be output based on the fuzzy intention and the second actual intention, further improving the usage experience of the voice interaction function.

[0183] In addition, by determining the application with the installed status as installed and the latest startup time among the multiple applications corresponding to the entity slot as the target entity, since the application with the latest startup time is the program that the user used last among the multiple applications corresponding to the entity slot, therefore, by determining this application as the target application, the confirmation efficiency of the user's true intention can be improved, thereby improving the response speed of voice data and further enhancing the usage experience of the voice interaction function.

[0184] It can be understood that in some other scenarios, when there is no ambiguous intention in the intention corresponding to the first voice data input by the user, there may be an entity slot corresponding to multiple applications in the slot corresponding to this intention. In this scenario, the user can be directly confirmed about their true intention regarding the multiple applications corresponding to the entity slot. Based on this, in some other embodiments, after S31, the voice interaction method may further include S35 to S38 as shown below, which are described in detail as follows: Figure 8 as shown below:

[0185] S35, when it is determined according to the information of the intention that there is no ambiguous intention in the intention and there is an entity slot corresponding to multiple applications in the slot corresponding to the intention, determine the target entity from the multiple applications corresponding to the entity slot.

[0186] Specifically, in combination with the foregoing embodiments, after the DM module receives the intention corresponding to the first voice data and the information of the intention, when it is determined according to the information of the intention that there is no ambiguous intention in the intention corresponding to the first voice data, but there is an entity slot corresponding to multiple applications in the slot corresponding to the intention, the DM module can determine the target entity from the multiple applications corresponding to the entity slot. It should be noted that the specific method for the DM module to determine the target application entity from the multiple applications corresponding to the entity slot can refer to the relevant descriptions in the foregoing embodiments and will not be elaborated here.

[0187] S36, output the fourth voice data.

[0188] Among them, the fourth voice data can be used to confirm the user's third actual intention regarding the target entity.

[0189] Exemplarily, the third actual intention includes a confirmation intention or a denial intention.

[0190] Optionally, when the third actual intention is a confirmation intention, it means that the user has confirmed the target entity, that is, it means that the target entity conforms to the user's true intention.

[0191] Optionally, when the third actual intention is a denial intention, it means that the user has denied the target entity, that is, it means that the target entity does not conform to the user's true intention.

[0192] Specifically, after the DM module determines the target entity, it can generate a third follow-up instruction based on the target entity and return the third follow-up instruction to the voice distribution module. The third follow-up instruction may carry third follow-up information for confirming whether the target entity conforms to the user's true intention, and the third follow-up instruction may be used to instruct the voice assistant to output the third follow-up information. The voice distribution module can transmit the third follow-up instruction to the voice assistant. After receiving the third follow-up instruction, the voice assistant can generate fourth voice data corresponding to the third follow-up information and broadcast the fourth voice data through the speaker. The third follow-up information can also be displayed on the screen.

[0193] Exemplarily, in combination with Figure 9 in (a), assume that the first voice data input by the user is "Open Ghost Blows Out the Light". The NLU module determines that the intention corresponding to the first voice data is "Open Ghost Blows Out the Light", and the confidence level of the intention "Open Ghost Blows Out the Light" is 1. The slot information includes: entity: video APP, reading APP. When the DM receives the intention of the first data and the information of the intention, it can first determine the target entity from "video APP" and "reading APP". Assume that the DM module determines "video APP" as the target entity. Then the DM module can generate third follow-up information "Confirm to open the movie of Ghost Blows Out the Light" for confirming whether the user's "video APP" conforms to the user's true intention based on the target entity, and send a third follow-up instruction carrying the third follow-up information "Confirm to open the movie of Ghost Blows Out the Light" to the voice distribution module. The voice distribution module can transmit the third follow-up instruction carrying the third follow-up information "Confirm to open the movie of Ghost Blows Out the Light" to the voice assistant. The voice assistant can generate fourth voice data with the content of "Confirm to open the movie of Ghost Blows Out the Light" according to the third follow-up information and broadcast the fourth voice data "Confirm to open the movie of Ghost Blows Out the Light" through the speaker. In addition, the voice assistant can also display the third follow-up information "Confirm to open the movie of Ghost Blows Out the Light" on the screen.

[0194] S37, receive the reply voice data for the fourth voice data, and process the reply voice data for the fourth voice data to obtain the third actual intention.

[0195] In this embodiment, after outputting the fourth voice data through the speaker, the voice assistant can control the microphone to collect the reply voice data for the fourth voice data, and can transmit the reply voice data for the fourth voice data to the voice processing application. The voice processing application can process the reply voice data for the fourth voice data to determine the third actual intention corresponding to the reply voice data for the fourth voice data.

[0196] Exemplarily, the voice assistant may determine the voice data collected within the first preset duration after outputting the fourth voice data as the reply voice data for the fourth voice data.

[0197] Exemplarily, in combination with Figure 9 in (b) thereof, in a specific implementation manner, the voice processing application may receive the reply voice data for the fourth voice data through the VAD module; after receiving the reply voice data for the fourth voice data, the VAD module may transmit the reply voice data for the fourth voice data to the voice distribution module. The voice distribution module may transmit the reply voice data for the fourth voice data to the ASR module. The ASR module may perform speech recognition on the reply voice data for the fourth voice data to obtain the reply text corresponding to the reply voice data for the fourth voice data, and return the reply text to the voice distribution module. The voice distribution module may transmit the reply text corresponding to the reply voice data for the fourth voice data to the NLU module. The NLU module may perform semantic recognition on the reply text to parse out the third actual intention corresponding to the reply voice data for the fourth voice data from the reply text.

[0198] S38. Output a response result for the first voice data according to the intention corresponding to the first voice data and the third actual intention.

[0199] Exemplarily, in combination with Figure 9 in (b) thereof, after the NLU module determines the third actual intention, it may return the third actual intention to the voice distribution module. The voice distribution module may transmit the third actual intention to the DM module. The DM module may determine whether to execute the target operation matching the intention corresponding to the first voice data according to the third actual intention.

[0200] In an alternative implementation manner, when the third actual intention is a confirmation intention, the DM module may determine to execute the target operation matching the intention corresponding to the first voice data. Based on this, S38 may specifically include: executing the target operation matching the intention corresponding to the first voice data according to the intention corresponding to the first voice data and the third actual intention, and outputting the execution result of the target operation.

[0201] Specifically, the DM module may execute the target operation by calling the target application corresponding to the target operation to obtain the execution result of the target operation, and return the execution result of the target operation to the voice distribution module. The voice distribution module may transmit the execution result of the target operation to the voice assistant. After receiving the execution result of the target operation, the voice assistant may broadcast the voice data corresponding to the execution result of the target operation through the speaker, and may also display the execution result of the target operation on the screen.

[0202] In another alternative implementation, when the third actual intention is a denial intention, the DM module may determine not to perform the target operation that matches the intention corresponding to the first voice data. Based on this, S38 may specifically include: outputting voice prompt data. The voice prompt data may be used to prompt not to perform the target operation that matches the intention corresponding to the first voice data.

[0203] Exemplarily, in combination with Figure 9 in (b), it is assumed that after the voice assistant outputs the fourth voice data "Confirm to turn on the Ghost Blowing Lamp movie?" through the speaker, the reply voice data "Cancel" for the fourth voice data is collected through the microphone. Then the voice assistant may transmit the reply voice data "Cancel" to the VAD module in the voice processing application. The VAD module may transmit the reply voice data "Cancel" to the voice distribution module. The voice distribution module may transmit the reply voice data "Cancel" to the ASR module. The ASR module may perform speech recognition on the reply voice data "Cancel" to obtain the reply text "Cancel" corresponding to the reply voice data "Cancel", and transmit the reply text "Cancel" to the voice distribution module. The voice distribution module may transmit the reply text "Cancel" to the NLU module. The NLU module may perform semantic recognition on the reply text "Cancel" and parse out that the third actual intention from the reply text "Cancel" is a denial intention. Based on this, the DM module may determine not to perform the operation of turning on the Ghost Blowing Lamp movie, generate a prompt instruction carrying the prompt information "Okay", and may return the prompt instruction to the voice distribution module. The voice distribution module may transmit the prompt instruction to the voice assistant. The voice assistant may generate the voice prompt data "Okay" corresponding to the prompt information based on the prompt instruction, and may broadcast the voice prompt data "Okay" through the speaker. In addition, the voice assistant may also display the prompt information "Okay" on the screen.

[0204] As can be seen from the above, in the voice interaction method provided in this embodiment, when there is no ambiguous intention in the intention corresponding to the first voice data, but there is an entity slot corresponding to multiple application programs in the slot corresponding to the intention, the fourth voice data for confirming the user's third actual intention for the target entity is output to further improve the user's true intention, so that a response result that better matches the user's true intention can be output based on the third intention, and the use experience of the voice interaction function is improved.

[0205] In a specific implementation, as Figure 10 shown, the DM module may specifically include a task orchestration unit, a pre-processing unit, and a task decision unit. In some embodiments, the task orchestration unit may be used to determine whether there is an ambiguous intention in the intention corresponding to the received first voice data according to the intention corresponding to the first voice data and the information of the intention.

[0206] Optionally, when there is an ambiguous intention in the intention corresponding to the first voice data, the task orchestration unit may instruct the preprocessing unit to generate a follow-up information for confirming the user's first actual intention for the ambiguous intention, and return a follow-up instruction carrying the follow-up information to the voice assistant, so as to instruct the voice assistant to output voice data for confirming the user's first actual intention for the ambiguous intention according to the follow-up information.

[0207] After that, after receiving the first actual intention returned by the voice distribution module, the task orchestration unit may directly transmit the first actual intention to the task decision unit, so that the task decision unit determines whether to execute the target operation matching the ambiguous intention according to the first actual intention. Optionally, when the task decision unit determines not to execute the target operation, it may generate a prompt message for not executing the target operation and return the prompt message to the voice assistant. Optionally, when the task decision unit determines to execute the target operation, it may execute the target operation matching the intention corresponding to the first voice data by calling the target application corresponding to the target operation and return the execution result of the target operation.

[0208] Optionally, when there is no ambiguous intention in the intention corresponding to the first voice data, the task orchestration unit may instruct the task decision unit to execute the target operation matching the intention corresponding to the first voice data and return the execution result of the target operation.

[0209] Based on the same technical concept, the embodiment of the present application also provides a computer-readable storage medium, which stores a computer-executable program. When the computer-executable program is called by a computer, the computer is enabled to execute one or more steps in any of the above method embodiments.

[0210] Based on the same technical concept, the embodiment of the present application also provides a chip system, including a processor, the processor is coupled with a memory, and the processor executes the computer-executable program stored in the memory to implement one or more steps in any of the above method embodiments. The chip system may be a single chip or a chip module composed of multiple chips.

[0211] Based on the same technical concept, the embodiment of the present application also provides a computer-executable program product. When the computer-executable program product runs on an electronic device, the electronic device is enabled to execute one or more steps in any of the above method embodiments.

[0212] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. It should be understood that the sequence numbers of the steps in the above embodiments do not indicate the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0213] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).

[0214] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by computer programs instructing relevant hardware. The programs can be stored in a computer-readable storage medium. When the programs are executed, they can include the processes of the above method embodiments. The foregoing storage media include: ROM or random access memory RAM, magnetic disks, or optical disks and other media that can store program codes.

[0215] The above is only the specific implementation manners of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of the present application should be covered by the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application should be subject to the protection scope of the claims.

Claims

1. A voice interaction method, characterized in that, Including: Receiving first voice data, processing the first voice data to obtain the intent corresponding to the first voice data and information of the intent; Outputting second voice data when it is determined that there is a fuzzy intent in the intent according to the information of the intent; The second voice data is used to confirm the user's first actual intent for the fuzzy intent; The first actual intent includes a confirmation intent or a denial intent; Receiving reply voice data for the second voice data, and processing the reply voice data to obtain the first actual intent; Outputting a response result for the first voice data according to the first actual intent.

2. The voice interaction method according to claim 1, wherein The information of the intent includes the quantity and confidence of the intent; correspondingly, when it is determined that there is a fuzzy intent in the intent according to the information of the intent, outputting second voice data includes: When the intent is a single intent and the confidence of the single intent is within a first confidence interval, determining that there is a fuzzy intent in the intent, and outputting second voice data based on the single intent and the information of the single intent.

3. The voice interaction method according to claim 1, characterized in that The information of the intent includes the quantity and confidence of the intent; correspondingly, when it is determined that there is a fuzzy intent in the intent according to the information of the intent, outputting second voice data includes: When the intent is multiple intents, determining that there is a fuzzy intent in the intent, and outputting second voice data based on the target intent with the highest confidence among the multiple intents and the information of the target intent.

4. The voice interaction method according to any one of claims 1-3, characterized in that, The information of the intent further includes slot information; correspondingly, outputting a response result for the first voice data according to the first actual intent includes: When the first actual intent is the confirmation intent, determining whether there is an entity slot corresponding to multiple application programs in the slot corresponding to the fuzzy intent according to the slot information of the fuzzy intent; When there is no such entity slot, performing a target operation matching the fuzzy intent according to the fuzzy intent and the slot information of the fuzzy intent, and outputting an execution result of the target operation.

5. The voice interaction method according to claim 4, characterized in that After determining whether there is an entity slot corresponding to multiple application programs in the slot corresponding to the fuzzy intent according to the slot information of the fuzzy intent, further including: When there is such entity slot, determining a target entity from the multiple application programs corresponding to the entity slot; Outputting third voice data; the third voice data is used to confirm the user's second actual intent for the target entity; the second actual intent includes a confirmation intent or a denial intent; Receiving reply voice data for the third voice data, and processing the reply voice data for the third voice data to obtain the second actual intent; Outputting a response result according to the fuzzy intent and the second actual intent.

6. The voice interaction method according to claim 5, wherein Outputting a response result according to the fuzzy intent and the second actual intent includes: When the second actual intent is the confirmation intent, performing a target operation matching the fuzzy intent according to the fuzzy intent and the second actual intent, and outputting an execution result of the target operation.

7. The voice interaction method according to claim 5, characterized in that Output a response result according to the fuzzy intention and the second actual intention, including: In the case where the second actual intention is a denial intention, output voice prompt data; the voice prompt data is used to prompt not to perform the target operation matching the fuzzy intention.

8. The voice interaction method according to claim 5, characterized in that Determine a target entity from multiple applications corresponding to the entity slot, including: Determine a target entity from multiple applications corresponding to the entity slot according to the installation status and the last startup time of the multiple applications corresponding to the entity slot.

9. The voice interaction method according to claim 8, wherein Determine a target entity from multiple applications corresponding to the entity slot according to the installation status and the last startup time of the multiple applications corresponding to the entity slot, including: Determine the application with the installed status of installed and the latest last startup time among the multiple applications corresponding to the entity slot as the target entity.

10. The voice interaction method according to any one of claims 1-9, characterized in that, Output a response result for the first voice data according to the first actual intention, including: In the case where the first actual intention is the denial intention, output voice prompt data; the voice prompt data is used to prompt not to perform the target operation matching the fuzzy intention.

11. The voice interaction method according to any one of claims 1-10, characterized in that, The information of the intention includes slot information; correspondingly, after obtaining the intention corresponding to the first voice data and the information of the intention, it further includes: In the case where it is determined according to the information of the intention that there is no fuzzy intention in the intention and there is an entity slot corresponding to multiple applications in the slot corresponding to the intention, determine a target entity from multiple applications corresponding to the entity slot; Output fourth voice data; the fourth voice data is used to confirm the user's third actual intention for the target entity; the third actual intention includes a confirmation intention or a denial intention; Receive the reply voice data for the fourth voice data, and process the reply voice data for the fourth voice data to obtain the third actual intention; Output a response result for the first voice data according to the intention corresponding to the first voice data and the third actual intention.

12. An electronic device, characterized in that, Including: One or more processors; One or more memories; The one or more memories store one or more computer-executable programs, and the one or more computer-executable programs include instructions. When the instructions are executed by the one or more processors, the electronic device executes the steps in the voice interaction method according to any one of claims 1-11.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer-executable program, and when the computer-executable program is called by an electronic device, the electronic device executes the steps in the voice interaction method according to any one of claims 1-11.

14. A chip system, characterized in that, The chip system is applied to an electronic device. The chip system includes a processor, and the processor is coupled to a memory. The memory is used to store computer program instructions. When the processor calls the computer program instructions, the electronic device implements the steps in the voice interaction method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Method and device for processing natural language dialogue, electronic device, and computer readable storage medium

    CN109002501A

  • Interaction device, interaction method, and program

    CN111489749A

  • Voice interaction method and system, storage medium and electronic equipment

    CN112331185A

  • Voice understanding method and device

    CN112740323A