Voice interaction method and related electronic equipment
Patent Information
- Application Number
- CN202380074611.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-04
- Filing Date
- 2023-09-07
- Publication Date
- 2025-05-30
AI Technical Summary
Existing voice assistants are easily awakened by mistake without the need for a wake-up word, resulting in reduced user experience and inability to effectively protect user privacy and security.
By receiving the voice signal and combining it with acceleration data, we use the voice detection model, pose detection model and audio-pose detection fusion model to calculate the confidence level to determine whether the voice signal is an instruction issued by the user, and then decide whether to wake up the voice assistant, and Perform voiceprint verification if necessary.
It effectively reduces the probability of false wake-up of the voice assistant, improves user experience, and protects user privacy and security.
Smart Images

Figure CN120077432A_ABST
Abstract
Description
Voice interaction method and related electronic equipment
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 4, 2022, with application number 202211376580.5 and invention name “A Voice Interaction Method and Related Electronic Devices”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of voice interaction, and in particular to a voice interaction method and related electronic equipment. Background Art
[0003] With the continuous development of smart electronic device technology, many electronic devices now feature voice assistants to facilitate user-to-device interaction. A voice assistant is an intelligent application that helps users solve problems through intelligent dialogue and real-time question-and-answer interaction. Generally, there are three different types of voice assistants: chat, question-and-answer, and command-based. Chat assistants are used for casual conversation and companionship, leveraging AI to engage in conversation and perceive user emotions. Question-and-answer assistants are used for knowledge acquisition, using dialogue to acquire knowledge or resolve questions. A common application is the intelligent customer service of various platforms. Command-based assistants are used for device control, using dialogue to control electronic devices and perform specific operations. Common applications include smart speakers and IoT devices. For example, voice control can say, "Turn on the air conditioner and set it to 25 degrees."
[0004] For some application scenarios where a wake-up word is not required to wake up the voice assistant, users can wake up the voice assistant without adding a specific wake-up word in the voice command, making the user's voice interaction with the electronic device more natural. In addition, not using a specific wake-up word when interacting with the electronic device is more in line with the user's habits.
[0005] Therefore, how to reduce the probability of voice assistants being woken up by mistake when voice interaction with voice assistants does not require wake-up words is an issue that technicians are increasingly concerned about.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide a voice interaction method and related electronic device method, which solve the problem of voice interaction applications being mistakenly awakened.
[0008] In a first aspect, an embodiment of the present application provides a voice interaction method, which is applied to an electronic device, wherein the electronic device includes a voice interaction application, and the method includes: receiving a first voice signal; when it is determined that the first voice signal is to be voice detected, obtaining voice signal data based on the first voice signal; processing the voice signal data through a voice detection model to obtain a first confidence level and voice data, the first confidence level being used to characterize the probability that the first voice signal is a voice command sent by a user to the electronic device; obtaining acceleration data of the electronic device, and obtaining posture information of the electronic device based on the acceleration data; processing the posture information through a posture detection model to obtain a second confidence level and target posture information, the second confidence level being used to characterize the probability that the electronic device is in a hand-held raised state; processing the target posture information and voice data through an audio-posture detection fusion model to obtain a third confidence level, the third confidence level being used to characterize the probability that the electronic device is in a hand-held raised state and the first voice signal is a voice command sent by the user to the electronic device; and determining whether to start the voice interaction application based on the first confidence level, the second confidence level, and the third confidence level.
[0009] In the above embodiment, after receiving a voice signal, if the electronic device determines that the voice signal requires voice detection, the electronic device processes the voice signal data of the voice signal through a voice detection model, processes the posture information through a posture detection module, and processes the high-order feature data output by the posture detection module and the voice detection model through an audio-posture monitoring model. These three models respectively output three confidence levels. Based on these three confidence levels, it is then determined whether the received voice signal is the target voice command for waking up the voice assistant. If so, the voice assistant is woken up; if not, the voice assistant is not woken up. Since the first confidence level is calculated by the voice detection model, the second confidence level is calculated by the posture detection model, and the third confidence level is calculated by the audio-posture detection fusion model, the first confidence level can exclude application scenarios where only the hand is raised, and the second confidence level can exclude application scenarios where only voice input is used. The third confidence level integrates the high-dimensional features of the voice information data and posture information, and can characterize the real-time correlation between the voice input and the posture state of the electronic device. Therefore, by using the above-mentioned first confidence level, second confidence level and third confidence level to determine whether the first voice signal is the target voice command, the obtained judgment result is more accurate, which can reduce the probability of the voice assistant being mistakenly awakened and improve the user experience.
[0010] In combination with the first aspect, in one possible implementation method, whether to start the voice interaction application is determined based on the first confidence level, the second confidence level and the third confidence level, specifically including: when the first confidence level is greater than or equal to the first confidence threshold, setting the first confidence level to 1; when the first confidence level is less than the first confidence threshold, setting the first confidence level to 0; when the second confidence level is greater than or equal to the second confidence threshold, setting the second confidence level to 1; when the second confidence level is less than the second confidence threshold, setting the second confidence level to 0; when the third confidence level is greater than or equal to the third confidence threshold, setting the third confidence level to 1; when the third confidence level is less than the third confidence threshold, setting the third confidence level to 0; performing a logical AND operation on the first confidence level, the second confidence level and the third confidence level to obtain a judgment result; and determining whether to start the voice interaction application based on the judgment result.
[0011] In this way, the electronic device can decide whether to send the first voice signal to the voice interaction application based on the judgment result, thereby avoiding the voice interaction application from being woken up by mistake and reducing the user's usage experience.
[0012] In combination with the first aspect, in one possible implementation, determining whether to start the voice interaction application is based on the judgment result, specifically including: when the judgment result is 1, starting the voice interaction application; when the judgment result is 0, not starting the voice interaction application.
[0013] In combination with the first aspect, in one possible implementation, the electronic device also includes a voiceprint detection module, which determines whether to start the voice interaction application based on the judgment result, specifically including: when the judgment result is 0, not starting the voice interaction application; when the judgment result is 1, the first voice signal is detected by the voiceprint detection module to see whether it is the voice of the target user, the target user being the user of the electronic device; if the judgment is yes, the voice interaction application is started; if the judgment is no, the voice interaction application is not started.
[0014] In this way, the voiceprint verification module sends the first voice signal to the voice assistant module only after determining that the first voice signal is a voice signal sent by the user. This method ensures that only the user of the electronic device can activate the voice assistant, ensuring the user's privacy and security while preventing the voice assistant from being accidentally triggered.
[0015] In combination with the first aspect, in one possible implementation method, determining whether to start a voice interaction application is based on the first confidence level, the second confidence level, and the third confidence level, specifically including: calculating a first weight value of the first confidence level, a second weight value of the second confidence level, and a third weight value of the third confidence level; calculating a fused confidence level based on the first confidence level, the first weight value, the second confidence level, the second weight value, the third confidence level, and the third weight value; and determining whether to start a voice interaction application based on the fused confidence level.
[0016] In conjunction with the first aspect, in one possible implementation, calculating a first weight value of the first confidence level, a second weight value of the second confidence level, and a third weight value of the third confidence level specifically includes: calculating according to the formula Calculate the first weight value, W1 is the first weight value, abs is the absolute value function, f m is the first confidence level of the speech detection model output this time, k is the number of the first Q confidence levels that are closest to the first confidence level of the current output; according to the formula Calculate the second weight value, W2 is the second weight value, L m is the second confidence level output by the posture detection model this time, k is the number of the first Q second confidence levels that are most adjacent to the second confidence level output this time; the third weight value is calculated according to the formula W3=1-W1-W2, where W3 is the third weight value.
[0017] In combination with the first aspect, in one possible implementation, based on the first confidence level, the first weight value, the second confidence level, the second weight value, the third confidence level, and the third weight value, the fused confidence level is calculated, specifically including: according to the formula K=f m W1+L m W2+R m W3 calculates the confidence after fusion; where K is the confidence after fusion, R m is the third confidence level.
[0018] In combination with the first aspect, in one possible implementation method, whether to start the voice interaction application is determined based on the fused confidence level, specifically including: if the fused confidence level is greater than or equal to the first start-up threshold, starting the voice interaction application; if the fused confidence level is less than the first start-up threshold, not starting the voice interaction application.
[0019] In combination with the first aspect, in one possible implementation, the electronic device includes a display screen. If the fused confidence level is less than a first start-up threshold and greater than or equal to a second start-up threshold, a prompt message is displayed on the display screen, and the prompt message is used to instruct the user to issue a voice command again; the second start-up threshold is less than the first start-up threshold.
[0020] In combination with the first aspect, in one possible implementation, the electronic device further includes a voiceprint detection module, which determines whether to start a voice interaction application based on the fused confidence level, specifically including: if the fused confidence level is less than a first start-up threshold, the voice interaction application is not started; if the fused confidence level is greater than or equal to the first start-up threshold, the first voice signal is detected by the voiceprint detection module to determine whether it is the voice of the target user, where the target user is the user of the electronic device; if the judgment is yes, the voice interaction application is started; if the judgment is no, the voice interaction application is not started.
[0021] In this way, the voiceprint verification module sends the first voice signal to the voice assistant module only after determining that the first voice signal is a voice signal sent by the user. This method ensures that only the user of the electronic device can activate the voice assistant, ensuring the user's privacy and security while preventing the voice assistant from being accidentally triggered.
[0022] In combination with the first aspect, in one possible implementation method, before obtaining voice signal data based on the first voice signal, it also includes: obtaining the signal strength value of the voice signal, the acceleration variance D1 of the electronic device on the x-axis, the acceleration variance D2 of the electronic device on the y-axis, and the acceleration variance D3 of the electronic device on the z-axis; and judging whether the first voice signal needs voice detection based on the signal strength value, D1, D2 and D3.
[0023] In this way, after receiving a voice signal, the electronic device first determines whether the voice signal requires voice detection through the wake-up-free first-level judgment module. For voice signals that do not require voice detection, the process is terminated and the voice signal is no longer processed. The voice signal is judged by the wake-up-free first-level judgment module, and most of the scenarios that are not intended by the user are filtered out, thereby avoiding the wake-up-free voice assistant in the electronic device and saving the computing resources of the electronic device.
[0024] In combination with the first aspect, in one possible implementation method, the speech data includes first speech data and second speech data, the first speech data is the high-order speech feature information output by the convolution layer of the speech detection model, and the second speech data is the high-order speech feature information output by the fully connected layer of the speech detection model; the target posture information includes first target posture information and second target posture information, the first target posture information is the high-order speech feature information output by the convolution layer of the posture detection model, and the second target posture information is the high-order speech feature information output by the fully connected layer of the posture detection model.
[0025] In a second aspect, an embodiment of the present application provides an electronic device, which includes: one or more processors, a display screen and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute: when it is determined that the first voice signal is to be subjected to voice detection, voice signal data is obtained based on the first voice signal; the voice signal data is processed by a voice detection model to obtain a first confidence level and voice data, the first confidence level is used to characterize the probability that the first voice signal is a voice instruction sent by the user to the electronic device; acceleration data of the electronic device is obtained, and posture information of the electronic device is obtained based on the acceleration data; the posture information is processed by a posture detection model to obtain a second confidence level and target posture information, the second confidence level is used to characterize the probability that the electronic device is in a hand-held raised state; the target posture information and voice data are processed by an audio-posture detection fusion model to obtain a third confidence level, the third confidence level is used to characterize the probability that the electronic device is in a hand-held raised state and the first voice signal is a voice instruction sent by the user to the electronic device; based on the first confidence level, the second confidence level and the third confidence level, it is determined whether to start a voice interaction application.
[0026] In combination with the second aspect, in one possible implementation method, the one or more processors call the computer instructions to cause the electronic device to execute: when the first confidence level is greater than or equal to the first confidence threshold, set the first confidence level to 1; when the first confidence level is less than the first confidence threshold, set the first confidence level to 0; when the second confidence level is greater than or equal to the second confidence threshold, set the second confidence level to 1; when the second confidence level is less than the second confidence threshold, set the second confidence level to 0; when the third confidence level is greater than or equal to the third confidence threshold, set the third confidence level to 1; when the third confidence level is less than the third confidence threshold, set the third confidence level to 0; perform a logical AND operation on the first confidence level, the second confidence level and the third confidence level to obtain a judgment result; and determine whether to start the voice interaction application based on the judgment result.
[0027] In combination with the second aspect, in one possible implementation, the one or more processors call the computer instruction to cause the electronic device to execute: when the judgment result is 1, start the voice interaction application; when the judgment result is 0, do not start the voice interaction application.
[0028] In combination with the second aspect, in one possible implementation method, the one or more processors call the computer instructions to cause the electronic device to execute: if the judgment result is 0, the voice interaction application is not started; if the judgment result is 1, the first voice signal is detected by the voiceprint detection module to see whether it is the voice of the target user, and the target user is the user of the electronic device; if the judgment is yes, the voice interaction application is started; if the judgment is no, the voice interaction application is not started.
[0029] In combination with the second aspect, in one possible implementation method, the one or more processors call the computer instructions to cause the electronic device to execute: calculate the first weight value of the first confidence level, the second weight value of the second confidence level, and the third weight value of the third confidence level; calculate the fused confidence level based on the first confidence level, the first weight value, the second confidence level, the second weight value, the third confidence level, and the third weight value; and determine whether to start the voice interaction application based on the fused confidence level.
[0030] In conjunction with the second aspect, in one possible implementation, the one or more processors call the computer instruction to cause the electronic device to execute: according to the formula Calculate the first weight value, W1 is the first weight value, abs is the absolute value function, f m is the first confidence level of the speech detection model output this time, k is the number of the first Q confidence levels that are closest to the first confidence level of the current output; according to the formula Calculate the second weight value, W2 is the second weight value, L m is the second confidence level output by the posture detection model this time, k is the number of the first Q second confidence levels that are most adjacent to the second confidence level output this time; the third weight value is calculated according to the formula W3=1-W1-W2, where W3 is the third weight value.
[0031] In conjunction with the second aspect, in one possible implementation, the one or more processors call the computer instruction to cause the electronic device to execute: according to the formula K=f m W1+L m W2+R m W3 calculates the confidence after fusion; where K is the confidence after fusion, R m is the third confidence level.
[0032] In combination with the second aspect, in one possible implementation, the one or more processors call the computer instructions to cause the electronic device to execute: if the fused confidence level is greater than or equal to a first start-up threshold, start the voice interaction application; if the fused confidence level is less than the first start-up threshold, do not start the voice interaction application.
[0033] In combination with the second aspect, in one possible implementation, the one or more processors call the computer instructions to cause the electronic device to execute: if the fused confidence level is less than the first start-up threshold and greater than or equal to the second start-up threshold, the control display screen displays a prompt message, and the prompt message is used to instruct the user to issue the voice command again; the second start-up threshold is less than the first start-up threshold.
[0034] In combination with the second aspect, in one possible implementation method, the one or more processors call the computer instructions to cause the electronic device to execute: if the fused confidence level is less than the first startup threshold, the voice interaction application is not started; if the fused confidence level is greater than or equal to the first startup threshold, the first voice signal is detected by the voiceprint detection module to see whether it is the voice of the target user, where the target user is the user of the electronic device; if the judgment is yes, the voice interaction application is started; if the judgment is no, the voice interaction application is not started.
[0035] In combination with the second aspect, in one possible implementation, the one or more processors call the computer instructions to cause the electronic device to execute: obtaining the signal strength value of the voice signal, the acceleration variance D1 of the electronic device on the x-axis, the acceleration variance D2 of the electronic device on the y-axis, and the acceleration variance D3 of the electronic device on the z-axis; and determining whether the first voice signal requires voice detection based on the signal strength value, D1, D2, and D3.
[0036] In a third aspect, an embodiment of the present application provides an electronic device comprising: a touch screen, a camera, one or more processors and one or more memories; the one or more processors are coupled to the touch screen, the camera, and the one or more memories, and the one or more memories are used to store computer program code, and the computer program code includes computer instructions. When the one or more processors execute the computer instructions, the electronic device executes the method described in the first aspect or any possible implementation method of the first aspect.
[0037] In a fourth aspect, an embodiment of the present application provides a chip system, which is applied to an electronic device, and the chip system includes one or more processors, which are used to call computer instructions to enable the electronic device to execute the method described in the first aspect or any possible implementation method of the first aspect.
[0038] In a fifth aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when run on an electronic device, enables the electronic device to execute the method described in the first aspect or any possible implementation of the first aspect.
[0039] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium comprising instructions, which, when executed on an electronic device, enables the electronic device to execute the method described in the first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figures 1A-1G are diagrams illustrating a set of scenarios of a voice interaction method provided in an embodiment of the present application;
[0041] FIG2 is a system framework diagram of a voice interaction method provided in an embodiment of the present application;
[0042] FIG3 is a flow chart of a voice interaction method provided in an embodiment of the present application;
[0043] FIG4 is an example diagram of a user interface provided in an embodiment of the present application;
[0044] FIG5A is a flow chart of another voice interaction method provided in an embodiment of the present application;
[0045] FIG5B is a structural diagram of a voiceprint detection model provided in an embodiment of the present application;
[0046] FIG6 is a schematic diagram of the hardware structure of the electronic device 100 provided in an embodiment of the present application;
[0047] FIG7 is a block diagram of the software structure of the electronic device 100 provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Mentioning "embodiment" in this article means that the specific features, structures or characteristics described in conjunction with the embodiment can be included in at least one embodiment of the present embodiment application. The appearance of this phrase in various positions in the specification does not necessarily mean the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It can be understood explicitly and implicitly by those skilled in the art that the embodiments described herein can be combined with other embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative work are within the scope of protection of this application.
[0049] In the specification, claims, and accompanying drawings of this application, the terms "first," "second," "third," and the like are used to distinguish different objects and are not used to describe a particular order. Furthermore, the terms "including," "comprising," "having," and any variations thereof are intended to cover non-exclusive inclusions. For example, a list of steps or elements may be included, or alternatively, steps or elements not listed may be included, or other steps or elements may be included that are inherent to the process, method, product, or apparatus.
[0050] Only part relevant to the present application is shown in the accompanying drawings, not all of it. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processing or methods depicted as flow charts. Although flow charts describe various operations (or steps) as sequential processing, many operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of various operations can be rearranged. When its operation is completed, the processing can be terminated, but can also have additional steps not included in the accompanying drawings. The processing can correspond to methods, functions, procedures, subroutines, subprograms, etc.
[0051] As used in this specification, the terms "component," "module," "system," "unit," and the like are used to refer to computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a unit can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or distributed between two or more computers. In addition, these units can be executed from various computer-readable media having various data structures stored thereon. Units can communicate, for example, through local and / or remote processes based on signals having one or more data packets (e.g., data from a second unit interacting with another unit in a local system, a distributed system, and / or a network. For example, the Internet interacts with other systems via signals).
[0052] In the embodiments of the present application, a voice interaction application is taken as an example to illustrate.
[0053] With the continuous development of intelligent electronic device technology, many electronic devices have the function of voice assistants to achieve the interaction between users and electronic devices. A voice assistant is an intelligent application that helps users solve problems through intelligent interaction such as intelligent conversation and instant Q&A. Generally, there are three different types of voice assistants: chatty, Q&A, and command. Chatty assistants are used to achieve the purpose of chatting and companionship, and communicate with users through AI technology to perceive users' emotions. Q&A assistants are used for knowledge acquisition, obtaining knowledge or solving doubts through conversations. A more common application is the intelligent customer service on various platforms. Command assistants are used for device control, controlling electronic devices through conversations to achieve certain operations. More common applications include smart speakers, IOT devices, etc. For example, voice control: "Turn on the air conditioner and set it to 25 degrees."
[0054] As shown in Figure 1A, when the user approaches the electronic device 100 and issues a voice command "Voice Assistant, open the music application" to the electronic device 100, in response to the user's voice command, the electronic device 100 opens the music application and displays the music interface as shown in Figure 1B. Or, as shown in Figure 1C, when the user approaches the electronic device 100 and issues a voice command "Voice Assistant, query the meaning of 'too highbrow to be popular'" to the electronic device 100, in response to the user's voice command, the electronic device 100 queries the meaning of "too highbrow to be popular" on the network and displays the query result on the user interface as shown in Figure 1D.
[0055] There are mainly two ways for users to wake up the voice assistant: One is that every time before waking up the voice assistant, the user needs to add a specific voice wake-up word in the voice command. The electronic device will wake up the voice assistant only when it detects the voice wake-up word in the user's voice command. Otherwise, the electronic device will not wake up the voice assistant. For electronic devices of different manufacturers, the wake-up words for waking up the voice assistant are different. For example, if the wake-up word for the electronic device of manufacturer 1 is "X Aitongxue". Then, when waking up the voice assistant in the electronic device of manufacturer 1, it is necessary to add "X Aitongxue" in front of the voice command. For example, X Aitongxue, please open the music application. This method of adding a wake-up word in front of the voice command often makes the voice interaction between the user and the electronic device unnatural and does not conform to the user's habits.
[0056] The other is that the user can wake up the voice assistant without adding a specific wake-up word in the voice command, that is: the user directly sends a voice command to the electronic device to wake up the voice assistant and instructs the voice assistant to perform corresponding operations. Exemplarily, as shown in Figure 1E, when the user approaches the electronic device 100 and issues a voice command "Open the music application" to the electronic device 100, in response to the user's voice command, the electronic device 100 opens the music application and displays the music interface as shown in Figure 1F.
[0057] In one possible implementation, a “voice-free wake-up” control may be included in the user interface for turning on the voice assistant function. As shown in Figure 1G, when the electronic device 100 detects an input operation (for example, a single click) for the “voice-free wake-up” control 101, in response to the operation, the electronic device 100 may turn on the “voice-free wake-up” function, that is, when the user sends a voice command to the electronic device 100, the voice assistant can be woken up without adding a specific wake-up word in the voice command. Optionally, after the electronic device 100 detects an input operation for the “voice-free wake-up” control 101, a prompt message may be displayed on the user interface as shown in Figure 1G to prompt the user to approach the microphone. For example, speak the command 2 to 5 cm close to the microphone at the bottom of the mobile phone.
[0058] With the second voice assistant wakeup method described above, users can wake up the voice assistant without adding a specific wakeup word to their voice commands, making voice interaction with electronic devices more natural. Furthermore, not using a specific wakeup word when interacting with electronic devices is more in line with user habits. However, the lack of a specific wakeup word to wake up the voice assistant can lead to false triggering of the voice assistant in electronic devices. For example, if a user places their phone on a table while searching for something and then asks someone where it is, without a wakeup word, the electronic device may activate the voice assistant to communicate with the user, resulting in a false triggering of the voice assistant. Alternatively, if a user places their phone on a table while speaking in a meeting, the electronic device may detect their voice signal and wake up the voice assistant, causing a false triggering of the voice assistant. Frequent false triggering of the voice assistant can cause inconvenience to the user and reduce the user experience.
[0059] Therefore, to address the above-mentioned issues, embodiments of the present application provide a method for voice interaction, comprising: an electronic device acquiring voice signal data and posture data; the voice signal data may include Mel-frequency cepstral coefficients of the voice signal received by multiple microphones of the electronic device and energy differences of the audio received by multiple microphones of the electronic device; and the posture data may include acceleration data in the x-axis direction, acceleration data in the y-axis direction, and acceleration data in the z-axis direction acquired by an accelerometer of the electronic device. The electronic device uses the voice signal data as input to a voice detection model, which processes the voice signal to obtain a first confidence level; the electronic device uses the posture data as input to the posture detection model, which processes the posture data and outputs a second confidence level; the electronic device uses the first voice data output by a convolutional layer of the voice detection model and the second voice data output by a fully connected layer of the voice detection model as input to a voice-pose detection model; and the electronic device uses the first target posture data output by the convolutional layer of the posture detection model and the second target posture data output by the fully connected layer of the posture detection model as input to the voice-pose detection model. The voice-position model processes the first voice data, the second voice data, the first target position data, and the second position data, and outputs a third confidence level. The electronic device determines whether to wake up the voice assistant based on the first confidence level, the second confidence level, and the third confidence level.
[0060] Below, the system framework of a voice interaction method provided by an embodiment of the present application is introduced in conjunction with Figure 2. As shown in Figure 2, the system architecture includes a wake-up-free judgment module and a voice assistant module. The wake-up-free judgment module is located in the digital audio processor layer (DSP layer), and the wake-up-free judgment module includes a wake-up-free first-level judgment module and a wake-up-free second-level judgment module. The voice assistant module is located in the application layer. After the wake-up-free judgment module receives the first voice signal, it first processes the first voice signal through the wake-up-free first-level judgment module to detect whether the first voice signal requires voice detection. If necessary, the first voice signal is sent to the wake-up-free second-level judgment module for voice detection. If the first voice signal is detected as a voice command sent to the electronic device, the wake-up-free judgment module sends the first voice signal to the voice assistant module, and then the voice assistant module performs the target operation according to the first voice signal.
[0061] Below, the process of a voice interaction method provided by an embodiment of the present application is introduced. Please refer to Figure 3, which is a flow chart of a voice interaction method provided by an embodiment of the present application. In Figure 3, the electronic device receives external voice signals through a microphone, and the number of microphones possessed by the electronic device is N, where N is an integer greater than or equal to 2. The electronic device shown in Figure 3 includes a wake-up-free judgment module and a voice assistant module. Among them, the wake-up-free judgment module includes a wake-up-free first-level judgment module and a wake-up-free second-level judgment module, and the wake-up-free second-level judgment module includes a voice detection model, a posture detection model, and a voice-posture detection model. For the sake of convenience, the embodiment of the present application is illustrated by taking N as 2. The specific process is as follows:
[0062] Step 301: The electronic device receives a first voice signal.
[0063] Specifically, the first voice signal may be a voice signal sent by a user, or a voice signal sent by other sound sources. An electronic device may have one or more microphones, and the electronic device may receive external voice signals through the microphones.
[0064] Step 302: The electronic device sends the first voice signal to the wake-up-free determination module.
[0065] Step 303: The wake-up-free judgment module processes the first voice signal through the wake-up-free first-level judgment module to obtain a first judgment result.
[0066] Specifically, after receiving a first voice signal, the electronic device may send the first voice signal to the wake-up-free judgment module. After receiving the first voice signal, the wake-up-free judgment module may process the first voice signal through the wake-up-free first-level judgment module. The wake-up-free first-level judgment module then calculates the signal strength of the first voice signal based on the received first voice signal to determine the strength of the first voice signal. If the first voice signal is weak, it is determined that the first voice signal is not a voice command issued to the electronic device. After calculating the signal strength of the first voice signal, the wake-up-free first-level judgment module may output a first judgment result. The first judgment result may be a first identifier or a second identifier. When the signal strength of the first voice signal is greater than or equal to a first threshold, the first judgment result is the first identifier, which indicates that the strength of the first voice signal is strong. When the signal strength of the first voice signal is less than the first threshold, the first judgment result is the second identifier, which indicates that the strength of the first voice signal is weak. The first threshold may be obtained based on historical values, empirical values, or experimental data, and is not limited in this embodiment of the present application.
[0067] Step 304: The wake-up-free judgment module processes the acceleration data through the wake-up-free first-level judgment module to obtain a second judgment result.
[0068] Specifically, after receiving the first voice signal, the electronic device can send the acceleration data to the wake-up-free judgment module. After receiving the acceleration data, the wake-up-free judgment module can process the acceleration data through the wake-up-free first-level judgment module to obtain a second judgment result. The acceleration data can be obtained by an acceleration sensor built into the electronic device, and the posture information can include the variance of the acceleration of the acceleration sensor on the x-axis, the variance of the acceleration on the y-axis, and the variance of the acceleration on the z-axis. Then, the electronic device judges whether the electronic device is in motion based on the variance of the acceleration corresponding to these three coordinate axes, thereby obtaining a second judgment result. The second judgment result includes a third identifier and a fourth identifier, the third identifier is used to indicate that the electronic device is in motion, and the fourth identifier is used to indicate that the electronic device is in a stationary state.
[0069] For example, the electronic device can determine whether the electronic device is in motion based on the variance of the acceleration corresponding to the three coordinate axes. The second judgment result can be obtained by: the electronic device can set variance thresholds for the three coordinate axes respectively, namely: a first variance threshold D1, a second variance threshold D2, and a third variance threshold D3. D1 corresponds to the x-axis, D2 corresponds to the y-axis, and D3 corresponds to the z-axis. The first variance threshold, the second variance threshold, and the third variance threshold can be the same or different, and can be obtained based on historical values, empirical values, or experimental data, and are not limited in the embodiments of the present application. If, among the variances of the accelerations corresponding to the three coordinate axes, any one has a variance greater than or equal to the corresponding variance threshold, the electronic device is determined to be in motion, and the second judgment result includes the first identifier. For example, if the variance of the acceleration corresponding to the x-axis is greater than or equal to D1, the electronic device is determined to be in motion. If the variances of the accelerations corresponding to the three coordinate axes are all less than the corresponding variance thresholds, the electronic device is determined not to be in motion.
[0070] In one possible implementation, if, among the variances of the accelerations corresponding to the three coordinate axes, only two of the variances of the accelerations are greater than or equal to the corresponding variance thresholds, the electronic device is determined to be in motion. For example, if the variance of the acceleration corresponding to the x-axis is greater than or equal to D1 and the variance of the acceleration corresponding to the y-axis is greater than or equal to D2, the electronic device is determined to be in motion. If, among the variances of the accelerations corresponding to the three coordinate axes, only one of the variances of the accelerations is greater than or equal to the corresponding variance thresholds, or if the variances of the accelerations corresponding to the three coordinate axes are all less than the corresponding variance thresholds, the electronic device is determined not to be in motion.
[0071] In one possible implementation, if all variances of the three accelerations corresponding to the three coordinate axes are greater than or equal to corresponding variance thresholds, the electronic device is determined to be in motion. Otherwise, the electronic device is determined not to be in motion.
[0072] It should be understood that step 303 can be executed before step 304, after step 304, or simultaneously with step 304. The embodiment of the present application does not limit the execution order of step 304 and step 303.
[0073] Step 305: The wake-up-free first-level judgment module determines whether to perform voice detection on the first voice signal according to the first judgment result and the second judgment result.
[0074] Specifically, after the wake-up-free first-level determination module calculates the first and second determination results, the electronic device can determine whether to perform voice detection on the first voice signal based on the first and second determination results, that is, whether the first voice signal is the target voice command for waking up the electronic device's voice assistant. If it is determined that voice detection is to be performed on the first voice signal, the electronic device executes step 306. If it is determined that voice detection is not to be performed on the first voice signal, the electronic device ends the process.
[0075] The method for the electronic device to determine whether to perform voice detection on the first voice signal may be: if the first judgment result includes the first identifier and the second judgment result includes the third identifier, the electronic device determines to perform voice detection on the first voice signal. Otherwise, the electronic device determines not to perform voice detection on the first voice signal.
[0076] For example, assuming that the first identifier and the third identifier are 1, and the second identifier and the fourth identifier are 0, the electronic device can perform a "logical AND" operation on the identifier in the first judgment result and the identifier in the second judgment result. If the operation result is 1, the electronic device determines to perform voice detection on the first voice signal. If the operation result is 0, the electronic device determines not to perform voice detection on the first voice signal.
[0077] Based on the signal strength of the first voice signal and the acceleration variance of the accelerometer, the electronic device can filter out most of the scenes that are not the user's intention. For example, the scene that is far away from the microphone of the electronic device (the strength of the voice signal received by the electronic device is weak), or the scene where the user chats while playing the electronic device (the variance of the acceleration data of the accelerometer is small), etc. is filtered out. For scenes that are the user's intention, the electronic device performs voice detection on the voice signal it receives, so as to make a more accurate judgment on whether the voice signal is an instruction to wake up the voice assistant. For scenes that are not the user's intention, the electronic device does not perform voice detection on the voice signal it receives and ends the process. Since voice detection on voice signals consumes a lot of computing resources. Therefore, before performing voice detection on the received voice signal, the electronic device determines whether the first voice signal meets the conditions for voice detection, which can greatly solve the computing resources of the electronic device and thus improve the working performance of the electronic device.
[0078] Step 306: The wake-up-free primary judgment module sends the first voice signal to the wake-up-free secondary judgment module.
[0079] Specifically, after the wake-up-free first-level judgment module determines to perform voice detection on the first voice signal, the wake-up-free first-level judgment module sends the first voice signal to the wake-up-free second-level judgment module so that the wake-up-free second-level judgment module performs voice detection on the first voice signal.
[0080] Step 307: The wake-up-free secondary module obtains voice signal data of the first voice signal.
[0081] Specifically, after receiving the first voice signal sent by the wake-up-free module and the wake-up-free module, the wake-up-free module processes the first voice signal to obtain voice signal data of the first voice signal.
[0082] The voice signal data may include the Mel-frequency cepstral coefficients of the first voice signal, and the energy difference M between the voice signal received by the first microphone and the first voice signal received by the second microphone. M is used to characterize the distance between the sound source (the sound source of the first voice signal) and the electronic device. The larger M is, the smaller the distance between the sound source and the electronic device is; the smaller M is, the larger the distance between the sound source and the electronic device is. The electronic device can set an energy threshold H. When M is greater than or equal to H, it can be considered that the sound source is close to the electronic device (for example, within 40 cm); when M is less than H, it can be considered that the sound source is far away from the electronic device (for example, beyond 40 cm). The Mel-frequency cepstral coefficient is a voice signal feature that conforms to the auditory characteristics of the human ear and captures more detailed features of the voice signal at low frequencies. In addition, when a user speaks to an electronic device at close range, there will be pop sounds at low frequencies. Therefore, the Mel-frequency cepstral coefficient is used as the input of the voice detection model, which can help the voice detection model extract the voice parameters of the first voice signal in the low-frequency domain.
[0083] Step 308: The wake-up-free secondary judgment module processes the voice signal data through a voice detection model to obtain a first confidence level, first voice data, and second voice data.
[0084] Specifically, the wake-up-free secondary determination module may process the voice signal data through a voice detection model to obtain a first confidence level, the first voice data, and the second voice data. The voice detection model may be a trained convolutional neural network, which may include a convolutional layer and a fully connected layer.
[0085] The wake-up-free secondary judgment module processes the voice signal data through the voice detection model. The convolution layer in the voice detection model first processes the voice signal to obtain and output the first voice data. The first voice data includes high-order feature information of Mel-frequency cepstral coefficients and high-order feature information of M. Then, the fully connected layer of the voice detection model processes the voice signal data processed by the convolution layer to obtain a first confidence level and second voice data. Among them, the second voice data includes high-order feature information of Mel-frequency cepstral coefficients and high-order feature information of M. The first confidence level is used to characterize the probability that the first voice signal is a voice command sent by the user to the electronic device.
[0086] Step 309: The wake-up-free secondary judgment module processes the posture information through the posture detection model to obtain the second confidence level, the first target posture information, and the second target posture information.
[0087] Optionally, before the wake-up-free secondary judgment module processes the posture information through the posture detection model, it can obtain acceleration data from the calculation sensor, and the acceleration data includes the acceleration data of the electronic device on the x-axis, the acceleration data on the y-axis, and the acceleration data on the z-axis. Then, based on the acceleration data of the electronic device on these three coordinate axes, the posture information of the electronic device is calculated. Among them, the posture information of the electronic device includes the absolute value of the acceleration data corresponding to the three coordinate axes of the x-axis, the y-axis, and the z-axis, and may also include the variance d1 of the acceleration data corresponding to the x-axis, the variance d2 of the acceleration data corresponding to the y-axis, and the variance d3 of the acceleration data corresponding to the z-axis. It may also include the mean p1 of the acceleration data corresponding to the x-axis, the mean p2 of the acceleration data corresponding to the y-axis, and the mean p3 of the acceleration data corresponding to the z-axis. It may also include the difference between d1 and p1, the difference between d2 and p2, and the difference between d3 and p3.
[0088] After obtaining the posture information, the wake-up-free secondary judgment module can detect the posture information through the posture detection model to determine whether the electronic device is currently in a hand-held raised state, and can also determine data such as the amplitude of the shaking of the electronic device in the hand-held raised state. Among them, the hand-held raised state can be understood as the user holding the electronic device in his hand. The electronic device can match the current application scenario with the first confidence level and the posture information, and determine whether the first voice signal is a voice command to wake up the voice assistant based on the application scenario. The electronic device can process the posture information through the posture detection model to obtain a second confidence level, first target posture information, and second target posture information.
[0089] The posture detection model can be a trained convolutional neural network model, which can include a convolution layer and a fully connected layer. Since the absolute values of the acceleration data corresponding to the three coordinate axes of x-axis, y-axis and z-axis, d1, d2 and d3 can represent whether the electronic device is in motion, p1, p2 and p3 can represent the amplitude of the movement of the electronic device, and the difference between d1 and p1, the difference between d2 and p2, and the difference between d3 and p3 can represent the motion state of the electronic device from other dimensions such as the smoothness of the movement. Therefore, the posture detection model can use the above-mentioned posture data to comprehensively judge whether the electronic device is in a handheld and raised state based on multiple aspects such as whether the electronic device is moving, the amplitude of the movement and the smoothness of the movement, thereby improving the accuracy of the posture detection model's judgment.
[0090] The convolutional layer in the pose detection model can first process the pose information and output first target pose information, which includes high-order feature information of the pose information. The fully connected layer of the pose detection model then processes the pose information processed by the convolutional layer to obtain a second confidence level and second target pose information. The second target pose information includes high-order feature information of the pose information, and the second confidence level is used to represent the probability that the electronic device is in the hand-held raised state.
[0091] It should be understood that step 308 can be executed before step 309, step 308 can also be executed after step 309, and step 308 can be executed simultaneously with step 309. The embodiment of the present application does not limit the execution order of step 308 and step 309.
[0092] Step 310: The wake-up-free secondary judgment module processes the first audio data, the second audio data, the first target posture information, and the second target posture information through the audio-posture detection fusion model to obtain a third confidence level.
[0093] Specifically, the audio-posture detection fusion model can be a trained convolutional neural network model, which is used to detect the probability that the first voice signal received by the electronic device is a voice command and the electronic device is currently in a hand-held raised state. After the electronic device processes the first audio data, the second audio data, the first target posture information, and the second target posture information through the audio-posture detection fusion model, it obtains a third confidence level. The third confidence level is used to characterize the probability that the first voice signal is a voice command and the electronic device is currently in a hand-held raised state, that is, to characterize the degree of match between the posture state of the electronic device and the voice signal received by the electronic device. The higher the third confidence level, the higher the probability that the first voice signal is a voice command and the electronic device is currently in a hand-held raised state, that is, the higher the real-time correlation between the presence of voice input in the electronic device and the electronic device being in a hand-held raised state.
[0094] Step 311: The wake-up-free secondary judgment module judges whether the first voice signal is a target voice command according to the first confidence level, the second confidence level, and the third confidence level.
[0095] Specifically, the target voice command is a command for waking up a voice assistant of the electronic device. If the electronic device determines that the first voice signal is the target voice command, step 312 is executed, otherwise, the process ends.
[0096] The electronic device may determine whether the first voice signal is a target voice command based on the first confidence level, the second confidence level, and the third confidence level in the following two methods:
[0097] The first method: Determine a first confidence flag based on a first confidence level, determine a second confidence flag based on a second confidence level, and determine a third confidence flag based on a third confidence level. When the first confidence level is greater than or equal to a first confidence threshold, the first confidence flag is 1; when the first confidence level is less than the first confidence threshold, the first confidence flag is 0. When the second confidence level is greater than or equal to the second confidence threshold, the second confidence flag is 1; when the second confidence level is less than the second confidence threshold, the second confidence flag is 0. When the third confidence level is greater than or equal to the third confidence threshold, the third confidence flag is 1; when the third confidence level is less than the third confidence threshold, the third confidence flag is 0. Then, the electronic device performs a "logical AND (&)" operation on the first confidence flag, the second confidence flag, and the third confidence flag to obtain a second judgment result. If the second judgment result is 1, the electronic device determines that the first voice signal is the target voice command; if the second judgment result is 0, the electronic device determines that the first voice signal is not the target voice command. The first confidence threshold, the second confidence threshold, and the third confidence threshold can be obtained from historical values, empirical values, or experimental data, and are not limited in this embodiment of the present application. Preferably, the first confidence threshold, the second confidence threshold, and the third confidence threshold can be 50%.
[0098] Second method: The electronic device may determine weighted values for the first, second, and third confidence levels using a formula. Then, based on the weighted values of the three confidence levels, the electronic device may fuse the three confidence levels to obtain a fused confidence level. The fused confidence level is then used to determine whether the first voice signal is the target voice command.
[0099] Exemplarily, the electronic device may calculate the weight value of the first confidence level using formula (1), which is as follows:
[0100] Among them, f m is the first confidence level output by the speech detection model this time, and k is the number of the first Q confidence levels adjacent to the first confidence level output by the speech detection model this time. For example, when k = 1, f k is the first confidence level of the last output of the speech detection model; when k=2, f k is the first confidence of the last output of the speech detection model... and so on. abs is the absolute value function.
[0101] The electronic device can calculate the weight value of the second confidence level by using formula (2), which is as follows:
[0102] Among them, L m is the second confidence level output by the posture detection model this time, and k is the number of the first Q second confidence levels adjacent to the second confidence level output by the posture detection model this time. For example, when k = 1, f k is the second confidence level of the last output of the pose detection model; when k=2, f k It is the second confidence of the last output of the pose detection model... and so on. abs is the absolute value function.
[0103] The electronic device can calculate the weight value of the third confidence level by using formula (3), which is as follows: W3 = 1 - W1 - W2 (3)
[0104] Then, the electronic device can calculate the fused confidence K according to formula (4), which is as follows: K = f m W1+L m W2+R m ·W3 (4)
[0105] Among them, K is the confidence after fusion, R mThis is the third confidence level of the current output of the audio-pose detection fusion model. After calculating K, the electronic device determines whether K is greater than or equal to the first activation threshold. If so, the electronic device determines that the first voice signal is the target voice command; otherwise, the electronic device determines that the first voice signal is not the target voice command. Preferably, the first activation threshold can be 60%.
[0106] Since the first confidence level is calculated by the voice detection model, the second confidence level is calculated by the posture detection model, and the third confidence level is calculated by the audio-posture detection fusion model, the first confidence level can exclude application scenarios where only the hand-held device is in the lifted state, and the second confidence level can exclude application scenarios where only voice input is used. The third confidence level integrates the high-dimensional features of voice information data and posture information, and can characterize the real-time correlation between voice input and posture state of electronic devices. Therefore, the judgment result obtained by judging whether the first voice signal is the target voice command through the above-mentioned first confidence level, second confidence level, and third confidence level is more accurate.
[0107] In one possible implementation, when it is determined by the second method above that the first voice signal is not the target voice command, the electronic device can also determine whether to display a prompt message based on the calculated and fused confidence level. If K is less than the first confidence threshold and greater than or equal to the second confidence threshold (the second confidence threshold is less than the first confidence threshold), the electronic device can display a prompt interface as shown in Figure 4 to prompt the user of problems that occurred when sending voice (for example, the voice is too small). In this way, the user knows where the problem is and makes timely improvements without waking up the voice assistant. Among them, the first confidence threshold and the second confidence threshold can be obtained based on historical values, or based on empirical values, or based on experimental data, and the embodiments of the present application are not limited thereto. Preferably, the second startup threshold can be 50%.
[0108] Step 312: The wake-up-free secondary judgment module sends the first voice signal to the voice assistant module.
[0109] Step 313: The voice assistant module parses the first voice signal and performs a first operation according to the first voice signal.
[0110] Specifically, after the wake-up-free secondary judgment module sends the first voice signal to the voice assistant module, the voice assistant module receives and analyzes the first voice signal, thereby obtaining the operation instruction, and performs the first operation according to the operation instruction.
[0111] For example, if a user sends a voice message to the electronic device saying "Open the camera app, I want to take a photo," the voice assistant module analyzes the first voice signal corresponding to the voice message and extracts the instruction "Open the camera app." Therefore, the voice assistant module can launch the camera app based on the instruction. The operation of launching the camera app by the voice assistant module is the first operation.
[0112] In an embodiment of the present application, after receiving a voice signal, the electronic device first determines whether the voice signal requires voice detection through the wake-up-free first-level judgment module. For voice signals that do not require voice detection, the process ends and the voice signal is no longer processed. The first-level judgment module judges the voice signal, filtering out most of the scenarios that are not intended by the user, thereby avoiding the voice assistant in the electronic device from being woken up, and saving the computing resources of the electronic device. If it is determined that the voice signal requires voice detection, the electronic device processes the voice signal data of the voice signal through the voice detection model, processes the posture information through the posture detection module, and processes the high-order feature data output by the posture detection module and the voice detection model through the audio-posture monitoring model. These three models output three confidence levels respectively, and then determine whether the received voice signal is the target voice command to wake up the voice assistant based on these three confidence levels. If so, the voice assistant is woken up, and if not, the voice assistant is not woken up. Since the first confidence level is calculated by the voice detection model, the second confidence level is calculated by the posture detection model, and the third confidence level is calculated by the audio-posture detection fusion model. The first confidence level can exclude scenarios where only the device is held up, the second confidence level can exclude scenarios where only voice input is used, and the third confidence level integrates the high-dimensional features of voice information data and posture information to characterize the real-time correlation between voice input and posture state of the electronic device. Therefore, using the first, second, and third confidence levels to determine whether the first voice signal is the target voice command yields a more accurate result, reducing the probability of the voice assistant being mistakenly awakened and improving the user experience.
[0113] In the embodiment of Figure 3 above, the process of a voice interaction method provided by an embodiment of the present application is introduced. Below, in conjunction with the accompanying drawings, another voice interaction method provided by an embodiment of the present application is introduced. In this method, after the wake-up-free judgment module determines that the first voice signal is the target voice command, the wake-up-free judgment module sends the first voice signal to the voiceprint verification module. After the voiceprint verification module determines that the first voice signal is a voice signal sent by the user himself, the first voice signal is sent to the voice assistant module. Through this method, only the user of the electronic device can wake up the voice assistant, which ensures the privacy and security of the user while ensuring that the voice assistant is not triggered by mistake.
[0114] Next, another voice interaction method proposed in the embodiment of the present application will be introduced in conjunction with FIG5A. Please refer to FIG5A, which is a flow chart of another voice interaction method provided in the embodiment of the present application. The specific process is as follows:
[0115] Step 501: The electronic device receives a first voice signal.
[0116] Step 502: The electronic device sends a first voice signal to the wake-up-free determination module.
[0117] Step 503: The wake-up-free judgment module processes the first voice signal through the wake-up-free first-level judgment module to obtain a first judgment result.
[0118] Step 504: The wake-up judgment module processes the acceleration data through the wake-up-free first-level judgment module to obtain a second judgment result.
[0119] Step 505: The wake-up-free first-level judgment module determines whether to perform voice detection on the first voice signal according to the first judgment result and the second judgment result.
[0120] Step 506: The wake-up-free primary determination module sends the first voice signal to the wake-up-free secondary determination module.
[0121] Step 507: The wake-up-free secondary module obtains voice signal data of the first voice signal.
[0122] Step 508: The wake-up-free secondary judgment module processes the voice signal data through a voice detection model to obtain a first confidence level, first voice data, and second voice data.
[0123] Step 509: The wake-up-free secondary judgment module processes the posture information through the posture detection model to obtain the second confidence level, the first target posture information, and the second target posture information.
[0124] Step 510: The wake-up-free secondary judgment module processes the first audio data, the second audio data, the first target posture information, and the second target posture information through the audio-posture detection fusion model to obtain a third confidence level.
[0125] Step 511: The wake-up-free secondary judgment module judges whether the first voice signal is a target voice command according to the first confidence level, the second confidence level, and the third confidence level.
[0126] If yes, execute step 512; if no, end the process.
[0127] Steps 501 to 511 may refer to steps 301 to 311 in the embodiment of FIG. 3 , and will not be described in detail here.
[0128] Step 512: The wake-up-free secondary judgment module sends the first voice signal to the voiceprint verification module.
[0129] Step 513: The voiceprint verification module verifies whether the first voice signal is a voice signal emitted by the user of the electronic device.
[0130] Specifically, the voiceprint verification module can be a pre-trained neural network model. As shown in Figure 5B, the user can enter a registration voice message according to the electronic device's prompts, for example, by saying "I look great today" or "Play today's news." The electronic device can extract voice feature information (e.g., frequency, loudness, pitch, timbre, etc.) from the user's registration voice message and use this extracted voice feature information as input to the acoustic model. The acoustic model processes this voice feature information and outputs the user's voiceprint feature information. This voiceprint feature information is then input to the back-end decision module, which processes the voiceprint feature information and outputs a difference function. This difference function measures the degree of difference between the voiceprint feature information output by the acoustic model and the user's actual voiceprint feature information. The larger the difference function, the greater the difference, and the smaller the difference function, the smaller the difference. The electronic device then adjusts the network structure or parameters of the acoustic model based on the difference function, so that the voiceprint feature information output by the acoustic model closely matches the user's voiceprint feature information. The voiceprint feature information is used to characterize the elements of the user's voice, and may include the pitch and timbre of the user's voice, as well as the loudness of the user's voice.
[0131] When the voiceprint verification module receives the first voice signal (input voice), it extracts voice feature information from the first voice signal and uses it as input to the acoustic model. The acoustic model processes the voice feature information and outputs voiceprint feature information corresponding to the first voice signal. This voiceprint feature information is then input to the backend decision module, which determines whether the voiceprint feature information is consistent with the user's voiceprint feature information. If so, the process proceeds to step 515. If not, the process ends.
[0132] Step 514: The voiceprint verification module sends the first voice signal to the voice assistant module.
[0133] Step 515: The voice assistant module parses the first voice signal and performs a first operation according to the first voice signal.
[0134] For step 515, reference may be made to step 313 in the embodiment of FIG. 3 , which will not be described in detail here.
[0135] It should be noted that, for simplicity of description, the above method embodiments are described as a series of actions. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described. Furthermore, those skilled in the art should also be aware that the embodiments described in this specification are preferred embodiments, and the actions involved are not necessarily required by the present invention.
[0136] The structure of the electronic device 100 is introduced below. Please refer to Figure 6, which is a schematic diagram of the hardware structure of the electronic device 100 provided in an embodiment of the present application.
[0137] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0138] It should be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in FIG6, or may combine or separate certain components, or may have different component arrangements. The components shown in FIG6 may be implemented in hardware, software, or a combination of software and hardware.
[0139] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0140] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0141] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0142] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.
[0143] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as Wi-Fi networks), Bluetooth (BT), BLE broadcast, global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0144] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0145] Display screen 194 is used to display images, videos, and the like. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, electronic device 100 may include one or N display screens 194, where N is a positive integer greater than one.
[0146] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.
[0147] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.
[0148] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0149] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU can enable intelligent cognitive applications in electronic device 100, such as image recognition, face recognition, speech recognition, and text comprehension.
[0150] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0151] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.
[0152] The speaker 170A, also called a "speaker", is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or listen to hands-free calls through the speaker 170A.
[0153] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or a voice message, the user can place the receiver 170B close to the ear to hear the voice.
[0154] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to realize sound signal collection, noise reduction, and identification of sound sources, and realize directional recording function, etc.
[0155] The pressure sensor 180A is used to sense the pressure signal and convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194 .
[0156] The air pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates the altitude using the air pressure value measured by the air pressure sensor 180C to assist in positioning and navigation.
[0157] The magnetic sensor 180D includes a Hall sensor, and the electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip leather case.
[0158] Accelerometer 180E can detect the magnitude of acceleration of electronic device 100 in all directions (generally three axes). It can also detect the magnitude and direction of gravity when electronic device 100 is stationary. It can also be used to identify the electronic device's posture, enabling applications such as switching between landscape and portrait modes and pedometers.
[0159] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to implement fingerprint unlocking, access application locks, fingerprint photography, fingerprint call answering, etc.
[0160] The touch sensor 180K is also called a "touch panel." The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also called a "touch screen." The touch sensor 180K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operations can be provided via the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, in a location different from that of the display screen 194.
[0161] The bone conduction sensor 180M can obtain a vibration signal. In some embodiments, the bone conduction sensor 180M can obtain a vibration signal of a vibrating bone mass in a human vocal part.
[0162] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservice architecture, or a cloud architecture. The embodiment of the present invention takes the Android system with a layered architecture as an example to exemplify the software structure of the electronic device 100. Figure 7 is a software structure block diagram of the electronic device 100 of an embodiment of the present application. The layered architecture divides the software into several layers, each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely, the application layer, the application framework layer, the hardware abstraction layer (HAL layer), the kernel layer, and the digital signal processing layer.
[0163] The application layer can include a series of application packages. As shown in Figure 7, the application package can include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, voice assistant module, video, etc.
[0164] The voice assistant module is used to parse the user's voice commands and perform relevant operations according to the user's voice commands, thereby realizing voice interaction between the electronic device and the user.
[0165] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions. As shown in Figure 7, the application framework layer may include a window manager, content provider, view system, telephony manager, resource manager, notification manager, and so on.
[0166] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.
[0167] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.
[0168] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.
[0169] The phone manager is used to provide communication functions of the electronic device 100, such as management of call status (including answering, hanging up, etc.).
[0170] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0171] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically without user interaction. For example, the Notification Manager is used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the top status bar of the system as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include text messages in the status bar, beeps, vibrations on electronic devices, and flashing indicator lights.
[0172] The hardware abstraction layer includes a voiceprint verification module, which is used to determine whether a received voice signal is a voice signal sent by a user.
[0173] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.
[0174] The digital signal processing layer includes a wake-up-free judgment module, which is used to judge whether the received voice signal is a voice signal to wake up the voice assistant in the electronic device.
[0175] It should be noted that, for the above method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the order of the actions described. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the present invention. The embodiments of the present application can be combined arbitrarily to achieve different technical effects.
[0176] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described herein are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive).
[0177] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium. When executed, the program can include the processes in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0178] In short, the above description is only an embodiment of the technical solution of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made based on the disclosure of the present invention should be included in the scope of protection of the present invention.
Claims
1. A voice interaction method, characterized in that: Applied to an electronic device, the electronic device including a voice interaction application, the method comprising: receiving a first voice signal; In a case where it is determined that the first speech signal is to be subjected to speech detection, obtaining speech signal data based on the first speech signal; Processing the voice signal data through a voice detection model to obtain a first confidence level and voice data, wherein the first confidence level is used to represent a probability that the first voice signal is a voice command sent by a user to the electronic device; Acquiring acceleration data of the electronic device, and obtaining posture information of the electronic device based on the acceleration data; Processing the posture information through a posture detection model to obtain a second confidence level and target posture information, wherein the second confidence level is used to represent a probability that the electronic device is in a hand-held raised state; Processing the target posture information and the voice data through an audio-posture detection fusion model to obtain a third confidence level, where the third confidence level is used to represent a probability that the electronic device is in a hand-held raised state and the first voice signal is a voice command sent to the electronic device by a user; Whether to start the voice interaction application is determined based on the first confidence level, the second confidence level, and the third confidence level.
2. The method according to claim 1, wherein The determining whether to start the voice interaction application based on the first confidence level, the second confidence level, and the third confidence level specifically includes: When the first confidence level is greater than or equal to a first confidence threshold, setting a first confidence flag to 1; When the first confidence level is less than a first confidence threshold, setting the first confidence flag to 0; When the second confidence level is greater than or equal to a second confidence threshold, setting the second confidence flag to 1; When the second confidence level is less than a second confidence threshold, setting the second confidence flag to 0; When the third confidence level is greater than or equal to a third confidence threshold, setting a third confidence flag to 1; When the third confidence level is less than a third confidence threshold, setting the third confidence flag to 0; Performing a logical AND operation on the first confidence identifier, the second confidence identifier, and the third confidence identifier to obtain a judgment result; Determine whether to start the voice interaction application based on the judgment result.
3. The method according to claim 2, wherein The determining whether to start the voice interaction application according to the judgment result specifically includes: When the judgment result is 1, starting the voice interaction application; When the judgment result is 0, the voice interaction application is not started.
4. The method according to claim 2, wherein The electronic device further includes a voiceprint detection module, and the determining whether to start the voice interaction application according to the judgment result specifically includes: If the judgment result is 0, the voice interaction application is not started; If the judgment result is 1, the first voice signal is detected by a voiceprint detection module to determine whether it is the voice of a target user, where the target user is the user of the electronic device; If the judgment is yes, start the voice interaction application; If the judgment is no, the voice interaction application is not started.
5. The method according to claim 1, wherein The determining whether to start the voice interaction application based on the first confidence level, the second confidence level, and the third confidence level specifically includes: Calculating a first weight value of the first confidence level, a second weight value of the second confidence level, and a third weight value of the third confidence level; Calculating a fused confidence level based on the first confidence level, the first weight value, the second confidence level, the second weight value, the third confidence level, and the third weight value; Based on the fused confidence level, it is determined whether to start the voice interaction application.
6. The method according to claim 5, wherein The calculating of the first weight value of the first confidence level, the second weight value of the second confidence level, and the third weight value of the third confidence level specifically includes: According to the formula Calculate the first weight value, W1 is the first weight value, abs is the absolute value function, f m is the first confidence level output by the speech detection model this time, and k is the number of the first Q first confidence levels that are most adjacent to the first confidence level output this time; According to the formula Calculate the second weight value, W2 is the second weight value, L m is the second confidence level output by the pose detection model this time, and k is the number of the first Q second confidence levels that are most adjacent to the second confidence level output this time; The third weight value is calculated according to the formula W3=1-W1-W2, where W3 is the third weight value.
7. The method according to claim 6, wherein The calculating a fused confidence level based on the first confidence level, the first weight value, the second confidence level, the second weight value, the third confidence level, and the third weight value specifically includes: According to the formula K = f m W1+L m W2+R m W3 calculates the confidence level after fusion; Wherein, K is the confidence after fusion, and R m is the third confidence level.
8. The method according to any one of claims 5 to 7, wherein: The determining whether to start the voice interaction application based on the fused confidence level specifically includes: If the fused confidence level is greater than or equal to a first activation threshold, activating the voice interaction application; If the fused confidence level is less than the first activation threshold, the voice interaction application is not activated.
9. The method according to claim 8, wherein The electronic device includes a display screen, and if the fused confidence level is less than a first start-up threshold and greater than or equal to a second start-up threshold, a prompt message is displayed on the display screen, wherein the prompt message is used to instruct the user to issue a voice command again; the second start-up threshold is less than the first start-up threshold.
10. The method according to any one of claims 5 to 7, characterized in that The electronic device further includes a voiceprint detection module, and the determining whether to start the voice interaction application based on the fused confidence level specifically includes: If the fused confidence level is less than a first activation threshold, not activating the voice interaction application; If the fused confidence level is greater than or equal to a first start threshold, detecting whether the first voice signal is the voice of a target user through a voiceprint detection module, where the target user is the user of the electronic device; If the judgment is yes, start the voice interaction application; If the judgment is no, the voice interaction application is not started.
11. The method according to any one of claims 1 to 10, wherein: Before obtaining voice signal data based on the first voice signal, the method further includes: Acquire a signal strength value of the voice signal, an acceleration variance D1 of the electronic device on the x-axis, an acceleration variance D2 of the electronic device on the y-axis, and an acceleration variance D3 of the electronic device on the z-axis; It is determined whether voice detection is required for the first voice signal based on the signal strength value, D1, D2, and D3.
12. The method according to any one of claims 1 to 11, wherein: The speech data includes first speech data and second speech data, the first speech data being high-order speech feature information output by a convolutional layer of the speech detection model, and the second speech data being high-order speech feature information output by a fully connected layer of the speech detection model; The target posture information includes first target posture information and second target posture information, the first target posture information is the high-order speech feature information output by the convolutional layer of the posture detection model, and the second target posture information is the high-order speech feature information output by the fully connected layer of the posture detection model.
13. An electronic device, characterized in that: include: Memory, processor, and touch screen; including: The touch screen is used to display content; The memory is used to store a computer program, wherein the computer program includes program instructions; The processor is configured to call the program instructions so that the electronic device executes the method according to any one of claims 1 to 12.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.