Electronic device and its wake-up method
By calculating the confidence of wake-up words and sound source orientation, combined with confidence threshold screening, the problem of false wake-up of electronic devices is solved and the user experience is improved.
Patent Information
- Application Number
- CN202011063583.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-30
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-09-30
AI Technical Summary
Existing electronic devices are prone to false wake-up when recognizing wake-up words, resulting in a decline in user experience.
By receiving sound, the confidence of wake-up words is calculated, and the sound source orientation and confidence threshold are used for screening to reduce the probability of false wake-up.
It effectively reduces the false wake-up rate of electronic devices and improves user experience.
Smart Images

Figure CN114360546B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and particularly relates to an electronic device and a wake-up method thereof. Background Art
[0002] An electronic device can perform functions through voice interaction with a user. Such an electronic device includes a sound pickup (such as a microphone array) and a speaker (such as a loudspeaker), and has a sound pickup function and a playback function. For example, a smart speaker, a smart phone, a smart TV, etc. Before a user interacts with an electronic device by voice, the electronic device needs to be woken up. Through waking up, the electronic device can enter the working state from the standby state. Generally speaking, an electronic device determines whether to wake up by identifying whether the received sound contains a preset wake-up word.
[0003] Taking the electronic device as a smart speaker and the wake-up word of the smart speaker being "Xiaoyi Xiaoyi" as an example. If a user emits a sound containing "Xiaoyi Xiaoyi", and the smart speaker detects "Xiaoyi Xiaoyi" from the received sound, the smart speaker wakes up. Sometimes, the smart speaker can also play a wake-up response voice to interact with the user by voice. For example, "Xiaoyi here. What can I do for you". However, in some scenarios, a user or other device emits a sound, but the emitted sound does not contain "Xiaoyi Xiaoyi", yet the smart speaker is woken up by mistake. For example, the user is watching TV. The user does not emit "Xiaoyi Xiaoyi", and the sound emitted by the TV does not contain "Xiaoyi Xiaoyi" either, but the smart speaker is woken up by mistake. In this way, it affects the normal life of the user, and the user also needs to turn off the smart speaker additionally, bringing a bad experience to the user. Summary of the Invention
[0004] In order to solve the above technical problems existing in the prior art, this application provides an electronic device and a wake-up method thereof, which can reduce the false wake-up rate of the electronic device and improve the user experience.
[0005] In a first aspect, a wake-up method is provided. The method is applied to an electronic device including a pickup and a speaker, and the pickup includes a plurality of microphones. The method includes: receiving a sound; calculating a wake-word confidence level of the sound, where the wake-word confidence level is used to represent the probability that the sound includes a wake word; after the wake-word confidence level is greater than or equal to a first threshold, calculating the sound source orientation of the sound; after the sound source orientation matches a first orientation in a first orientation set, and after the first orientation confidence level corresponding to the matched first orientation is greater than or equal to a third threshold, waking up the electronic device; or, after the first orientation confidence level corresponding to the matched first orientation is less than the third threshold, not waking up the electronic device. Among them, the wake word is used to wake up the electronic device; the sound source orientation is the direction and position of the sound source relative to the electronic device; the first orientation set includes M first orientation elements, and each first orientation element includes a first orientation and a first orientation confidence level; the first orientation is the direction and position of the sound source that wakes up the electronic device relative to the electronic device, and is used to represent that the electronic device has been woken up in the first orientation; the first orientation confidence level is used to represent the probability of waking up the electronic device in the first orientation; M is a positive integer greater than or equal to 1. In this way, when the wake-word confidence level of the sound is greater than or equal to the first threshold, possible false awakenings are further filtered out according to the first orientation set, thereby reducing the false awakening probability of the electronic device and improving the user experience.
[0006] According to the first aspect, after the wake-word confidence level is greater than or equal to the first threshold, calculating the orientation of the sound source corresponding to the sound; including: after the wake-word confidence level is greater than or equal to the first threshold and less than a second threshold, calculating the orientation of the sound source corresponding to the sound. In this way, by setting the second threshold and filtering out the cases where the wake-word confidence level is between the first threshold and the second threshold, the processing efficiency of the electronic device can be improved.
[0007] According to the first aspect, or any implementation manner of the above first aspect, the sound source orientation matches a first orientation in the first orientation set; including: the angle deviation between the direction of the sound source orientation relative to the electronic device and the direction of a first orientation in the first orientation set relative to the electronic device is within a preset fourth threshold; and the position deviation between the position of the sound source orientation relative to the electronic device and the position of the first orientation relative to the electronic device is within a preset fifth threshold.
[0008] According to the first aspect, or any implementation of the above first aspect, after the sound source orientation does not match any of the first orientations in the first orientation set, a voiceprint is extracted from the sound; after the voiceprint matches a first voiceprint in the first voiceprint set, and the first voiceprint confidence corresponding to the first voiceprint is greater than or equal to a preset sixth threshold, the electronic device is awakened; or, after the first voiceprint confidence corresponding to the first voiceprint is less than the preset sixth threshold, the electronic device is not awakened. Wherein, the first voiceprint set includes L voiceprint elements, and each voiceprint element includes a first voiceprint and a first voiceprint confidence. The first voiceprint is used to represent the voiceprint for awakening the electronic device, and the first voiceprint confidence is used to represent the probability of the first voiceprint awakening the electronic device; L is a positive integer greater than or equal to 1. In this way, if possible false awakenings cannot be filtered out according to the first orientation set, possible false awakenings are further filtered out through the first voiceprint set, thereby reducing the false awakening probability of the electronic device and improving the user experience.
[0009] According to the first aspect, or any implementation of the above first aspect, after awakening the electronic device, the method further includes: updating the first orientation set and the first voiceprint set.
[0010] According to the first aspect, or any implementation of the above first aspect, after the wake-up word confidence is greater than or equal to the second threshold, the electronic device is awakened, and the first orientation set and the first voiceprint set are updated.
[0011] According to the first aspect, or any implementation of the above first aspect, after the wake-up word confidence is greater than or equal to the second threshold, the electronic device is awakened, and the first orientation set and the first voiceprint set are updated; including: after the wake-up word confidence is greater than or equal to the second threshold, the electronic device is awakened, and the first orientation set and the first voiceprint set are created; the orientation for awakening the electronic device is included in the first orientation set, the voiceprint for awakening the electronic device is included in the first voiceprint set, and an initial orientation confidence is given to the orientation included in the first orientation set, and an initial voiceprint confidence is given to the voiceprint included in the first voiceprint set.
[0012] Second aspect, a wake-up method is provided. The method is applied to an electronic device including a pickup and a speaker, and the pickup includes a plurality of microphones. The method includes: receiving a sound; calculating a wake-word confidence of the sound, where the wake-word confidence is used to represent the probability that the sound includes a wake word; after the wake-word confidence is greater than or equal to a first threshold, calculating a sound source orientation; after the sound source orientation matches one of the second orientations in a second orientation set, and after the second orientation confidence corresponding to the matched second orientation is greater than or equal to a seventh threshold, waking up the electronic device; or, if the second orientation confidence corresponding to the matched second orientation is less than the seventh threshold, then not waking up the electronic device. Wherein, the wake word is used to wake up the electronic device; the sound source orientation is the direction and position of the sound source relative to the electronic device; the second orientation set includes N second orientation elements, and each second orientation element includes a second orientation and a second orientation confidence. The second orientation is the direction and position of the sound source that does not wake up the electronic device relative to the electronic device, and is used to indicate that the electronic device is not woken up in the second orientation. The second orientation confidence is used to represent the probability that the electronic device is not woken up in the second orientation; N is a positive integer greater than or equal to 1. In this way, when the wake-word confidence of the sound is greater than or equal to the first threshold, possible false wake-ups are further filtered out according to the second orientation set, thereby reducing the false wake-up probability of the electronic device and improving the user experience.
[0013] According to the second aspect, the sound source orientation matches one of the second orientations in the second orientation set, including: the angle deviation between the direction of the sound source orientation relative to the electronic device and the direction of one of the second orientations in the second orientation set relative to the electronic device is within a preset eighth threshold; and the position deviation between the position of the sound source orientation relative to the electronic device and the position of the second orientation relative to the electronic device is within a preset ninth threshold.
[0014] According to the second aspect, or any implementation manner of the above second aspect, the method further includes: after the sound source orientation does not match any of the second orientations in the second orientation set, extracting a voiceprint from the sound; after the voiceprint does not match any of the first voiceprints in a first voiceprint set, updating the second orientation set. Wherein, the first voiceprint set includes L voiceprint elements, and each voiceprint element includes a first voiceprint and a first voiceprint confidence. The first voiceprint is used to represent the voiceprint for waking up the electronic device, and the first voiceprint confidence is used to represent the probability that the first voiceprint wakes up the electronic device; L is a positive integer greater than or equal to 1. In this way, if possible false wake-ups cannot be filtered out according to the second orientation set, possible false wake-ups are further filtered out through the first voiceprint set, thereby reducing the false wake-up probability of the electronic device and improving the user experience.
[0015] According to the second aspect, or any implementation manner of the above second aspect, the method further includes: after the sound source orientation does not match any second orientation in the second orientation set, extracting a voiceprint from the sound; after the voiceprint matches a first voiceprint in the first voiceprint set, and after the first voiceprint confidence corresponding to the first voiceprint is greater than or equal to a preset tenth threshold, waking up the electronic device; or, after the first voiceprint confidence corresponding to the first voiceprint is less than the preset tenth threshold, not waking up the electronic device and updating the second orientation set. Wherein, the first voiceprint set includes L voiceprint elements, each voiceprint element includes a first voiceprint and a first voiceprint confidence, the first voiceprint confidence is used to represent the probability of the first voiceprint waking up the electronic device, and the first voiceprint is used to represent the voiceprint for waking up the electronic device; L is a positive integer greater than or equal to 1.
[0016] According to the second aspect, or any implementation manner of the above second aspect, after waking up the electronic device, the method further includes: updating the first voiceprint set; after not waking up the electronic device, the method further includes: updating the second orientation set.
[0017] According to the second aspect, or any implementation manner of the above second aspect, after the wake-up word confidence is greater than or equal to the second threshold, waking up the electronic device and updating the first voiceprint set. In this way, by setting the second threshold and screening out the cases where the wake-up word confidence is between the first threshold and the second threshold, the processing efficiency of the electronic device can be improved.
[0018] In a third aspect, a wake-up method is provided. The method is applied to an electronic device including a pickup and a speaker, and the pickup includes a plurality of microphones. The method includes: receiving a sound; calculating a wake-word confidence level of the sound, where the wake-word confidence level is used to represent the probability that the sound includes a wake word; after the wake-word confidence level is greater than or equal to a first threshold, calculating the sound source azimuth; after the sound source azimuth matches one of the second azimuths in the second azimuth set, and after the sound source azimuth does not match any of the first azimuths in the first azimuth set, and after the second azimuth confidence level corresponding to the matched second azimuth is greater than or equal to an eleventh threshold, waking up the electronic device; or, after the second azimuth confidence level corresponding to the matched second azimuth is less than the eleventh threshold, not waking up the electronic device. Wherein, the wake word is used to wake up the electronic device; the sound source azimuth is the direction and position of the sound source relative to the electronic device; the first azimuth set includes M first azimuth elements, and each first azimuth element includes a first azimuth and a first azimuth confidence level; the first azimuth is the direction and position of the sound source that wakes up the electronic device relative to the electronic device, and is used to represent that the electronic device has been woken up at the first azimuth; the first azimuth confidence level is used to represent the probability of waking up the electronic device at the first azimuth; the second azimuth set includes N second azimuth elements, and each second azimuth element includes a second azimuth and a second azimuth confidence level; the second azimuth is the direction and position of the sound source that has not woken up the electronic device relative to the electronic device, and is used to represent that the electronic device has not been woken up at the second azimuth; the second azimuth confidence level is used to represent the confidence level that the electronic device has not been woken up at the second azimuth; both M and N are positive integers greater than or equal to 1. In this way, when the wake-word confidence level of the sound is greater than or equal to the first threshold, possible false wake-ups are further filtered out according to the first azimuth set and the second azimuth set, thereby reducing the false wake-up probability of the electronic device and improving the user experience.
[0019] According to the third aspect, the sound source azimuth matching one of the second azimuths in the second azimuth set includes: the angle deviation between the direction of the sound source azimuth relative to the electronic device and the direction of one of the second azimuths relative to the electronic device is within a preset twelfth threshold; and, the position deviation between the position of the sound source azimuth relative to the electronic device and the position of the second azimuth relative to the electronic device is within a preset thirteenth threshold; the sound source azimuth not matching any of the first azimuths in the first azimuth set includes: the angle deviation between the direction of the sound source azimuth relative to the electronic device and the direction of any of the first azimuths relative to the electronic device is not within a preset fourteenth threshold; and, the position deviation between the position of the sound source azimuth relative to the electronic device and the position of any of the first azimuths relative to the electronic device is not within a preset fifteenth threshold.
[0020] According to a third aspect, or any implementation manner of the above third aspect, the method further includes: after the sound source orientation matches a first orientation in the first orientation set, and after the sound source orientation does not match any second orientation in the second orientation set, and after the first orientation confidence corresponding to the matched first orientation is greater than or equal to the sixteenth threshold, the electronic device is awakened; or, after the first orientation confidence corresponding to the matched first orientation is less than the sixteenth threshold, the electronic device is not awakened.
[0021] According to a third aspect, or any implementation manner of the above third aspect, the sound source orientation matches a first orientation in the first orientation set; including: the angle deviation between the direction of the sound source orientation relative to the electronic device and the direction of a first orientation in the first orientation set relative to the electronic device is within a preset fourteenth threshold; and the position deviation between the position of the sound source orientation relative to the electronic device and the position of the first orientation relative to the electronic device is within a preset fifteenth threshold; the sound source orientation does not match any second orientation in the second orientation set; including: the angle deviation between the direction of the sound source orientation relative to the electronic device and the direction of any second orientation in the second orientation set relative to the electronic device is not within a preset twelfth threshold; and the position deviation between the position of the sound source orientation relative to the electronic device and the position of any second orientation in the second orientation set relative to the electronic device is not within a preset thirteenth threshold.
[0022] According to a third aspect, or any implementation manner of the above third aspect, the method further includes: after the sound source orientation does not match any second orientation in the second orientation set, and after the sound source orientation does not match any first orientation in the first orientation set, the voiceprint is extracted from the sound; after the voiceprint matches a first voiceprint in the first voiceprint set, and after the first voiceprint confidence corresponding to the first voiceprint is greater than or equal to a preset sixteenth threshold, the electronic device is awakened, and the first orientation combination and the first voiceprint set are updated; or, after the first voiceprint confidence corresponding to the first voiceprint is less than the preset sixteenth threshold, the electronic device is not awakened, and the second orientation set is updated. Wherein, the first voiceprint set includes L voiceprint elements, each voiceprint element includes a first voiceprint and a first voiceprint confidence, the first voiceprint confidence is used to represent the probability of the first voiceprint awakening the electronic device, and the first voiceprint is used to represent the voiceprint for awakening the electronic device; L is a positive integer greater than or equal to 1. In this way, if possible false awakenings cannot be filtered out according to the first orientation set and the second orientation set, possible false awakenings are further filtered through the first voiceprint set, thereby reducing the false awakening probability of the electronic device and improving the user experience.
[0023] According to a third aspect, or any implementation of the above third aspect, the method further includes: after the voiceprint does not match any of the first voiceprints in the first voiceprint set, updating the second azimuth set.
[0024] According to a third aspect, or any implementation of the above third aspect, the method further includes: after waking up the electronic device, updating the first azimuth set; after not waking up the electronic device, updating the second azimuth set.
[0025] In a fourth aspect, an electronic device is provided. The electronic device includes a pickup and a speaker, the pickup includes a plurality of microphones, and the electronic device further includes: a processor; a memory; and a computer program, where the computer program is stored in the memory, and when the computer program is executed by the processor, the electronic device is caused to execute the methods described in the first aspect and any implementation of the first aspect, the second aspect and any implementation of the second aspect, and the third aspect and any implementation of the third aspect.
[0026] For the technical effects corresponding to the fourth aspect and any implementation of the fourth aspect, reference may be made to the technical effects corresponding to the first aspect and any implementation of the first aspect, the second aspect and any implementation of the second aspect, and the third aspect and any implementation of the third aspect, which will not be elaborated here.
[0027] In a fifth aspect, a computer-readable storage medium is provided. The computer-readable storage medium includes a computer program, and when the computer program runs on an electronic device, the electronic device is caused to execute the methods described in the first aspect and any implementation of the first aspect, the second aspect and any implementation of the second aspect, and the third aspect and any implementation of the third aspect, where the electronic device includes a pickup and a speaker, and the pickup includes a plurality of microphones.
[0028] For the technical effects corresponding to the fifth aspect and any implementation of the fifth aspect, reference may be made to the technical effects corresponding to the first aspect and any implementation of the first aspect, the second aspect and any implementation of the second aspect, and the third aspect and any implementation of the third aspect, which will not be elaborated here.
[0029] In a sixth aspect, a computer program product is provided. When it runs on a computer, the computer is caused to execute the methods described in the first aspect and any implementation of the first aspect, the second aspect and any implementation of the second aspect, and the third aspect and any implementation of the third aspect.
[0030] For the sixth aspect and the technical effects corresponding to any implementation manner in the sixth aspect, reference may be made to the above-mentioned first aspect and any implementation manner in the first aspect, the second aspect and any implementation manner in the second aspect, and the third aspect and the technical effects corresponding to any implementation manner in the third aspect, which will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application;
[0032] Figure 2 Schematic diagram of the software structure of the electronic device provided in the embodiment of the present application;
[0033] Figure 3 Schematic diagram of the scenario of the wake-up method provided in the embodiment of the present application;
[0034] Figure 4 Schematic diagram of the graphical user interface set by the user in the wake-up method provided in the embodiment of the present application;
[0035] Figure 5 Flowchart of an embodiment of the wake-up method provided in the embodiment of the present application;
[0036] Figure 6 Flowchart of another embodiment of the wake-up method provided in the embodiment of the present application;
[0037] Figure 7 Flowchart of yet another embodiment of the wake-up method provided in the embodiment of the present application;
[0038] Figure 8 Flowchart of yet another embodiment of the wake-up method provided in the embodiment of the present application;
[0039] Figure 9 Flowchart of yet another embodiment of the wake-up method provided in the embodiment of the present application;
[0040] Figure 10 Flowchart of yet another embodiment of the wake-up method provided in the embodiment of the present application;
[0041] Figure 11 Schematic diagram of the structural composition of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification and the appended claims of this application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, "at least one" and "one or more" mean one, two or more (including two). The term "and / or" is used to describe the relationship between associated objects and indicates that there can be three relationships; for example, A and / or B can mean: A exists alone, A and B exist simultaneously, or B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0043] Reference to "one embodiment" or "some embodiments" etc. described in this specification means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way. The term "connect" includes direct connection and indirect connection, unless otherwise stated.
[0044] The terms used in the embodiments section of this application are only used to explain the specific embodiments of this application and are not intended to limit this application.
[0045] In one example, the false wake-up probability generated by an electronic device is reduced by optimizing a preset wake-up word model in the electronic device. The main function of the wake-up word model is to detect a wake-up word from the sound picked up by the electronic device and obtain the probability that the sound contains the wake-up word. The wake-up word model is a trained machine learning model. For example, a model for detecting a wake-up word can be established in advance, and the model is trained with samples to obtain the wake-up word model. The above-mentioned pre-established model can be a neural network model, a Gaussian mixture model, a hidden Markov model, etc. The above-mentioned samples can be sounds containing wake-up words, or phoneme sequences of sounds containing wake-up words, or audio features of sounds containing wake-up words, etc. The sounds containing wake-up words can be recorded by different people in different scenarios. Using the sounds containing wake-up words recorded by different people in different scenarios can enable the trained wake-up word model to detect the wake-up words in the sounds in various scenarios. The sounds recorded in different scenarios do not only include wake-up words, but may include noise (such as non-wake-up words). In this way, if the sounds recorded in different scenarios are used as samples to train the wake-up word model, the wake-up word model will be contaminated, resulting in the wake-up word model possibly identifying the sounds including non-wake-up words as wake-up word sounds, thus causing false wake-up. Taking a smart speaker with the wake-up word "Xiaoyi Xiaoyi" as an example, after the wake-up word model trained based on the above method is set in the smart speaker, the wake-up word model may detect the sounds in the sound picked up by the smart speaker that are similar in pronunciation to "Xiaoyi Xiaoyi" or even completely different in pronunciation as sounds containing the wake-up word, thus causing the smart speaker to be falsely woken up.
[0046] In order to minimize the contamination of the wake-up word model caused by unclean samples and further avoid the false wake-up problem of the electronic device, it is necessary to continuously optimize the wake-up word model. Specifically, the wake-up word model is iteratively optimized through data annotation. However, data annotation requires manual annotation of the sounds used as samples, resulting in excessive consumption of human resources, and the optimized wake-up word model still has a certain probability of false wake-up problems. Therefore, the embodiments of the present application provide an electronic device and a wake-up method, which can reduce the false wake-up probability of the electronic device and improve the user experience.
[0047] The electronic device provided by the embodiments of the present application is an electronic device with a sound pickup function and an external voice function. For example: smart speakers, smartphones, tablets, personal computers (PCs), wearable devices (such as smart glasses, smart watches, smart bracelets, etc.), smart home appliances such as smart TVs, smart screens, intelligent connected vehicles (ICV), smart (motor) cars or in-vehicle devices, etc.
[0048] Exemplarily,Figure 1 A schematic structural diagram of the electronic device 100 is shown. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0049] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware. Exemplarily, the electronic device 100 may be a smart speaker. The smart speaker may include: a processor 110, an internal memory 121, a speaker 170A, and a microphone 170C.
[0050] The processor 110 may include one or more processing units. For example: the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc., in part or in whole. Among them, different processing units may be independent devices or integrated in one or more processors.
[0051] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, etc.
[0052] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple groups of I2S buses. The processor 110 can be coupled to the audio module 170 through the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit sound to the wireless communication module 160 through the I2S interface to implement the function of answering a phone call through a Bluetooth headset.
[0053] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled through the PCM bus interface. In some embodiments, the audio module 170 can also transmit sound to the wireless communication module 160 through the PCM interface to implement the function of answering a phone call through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0054] It can be understood that the interface connection relationships among the modules illustrated in the embodiments of the present invention are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0055] The electronic device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, which is connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0056] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0057] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.). In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.
[0058] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc. Such as music playback, recording, etc.
[0059] The audio module 170 is used to convert digital audio information into analog sound for output, and is also used to convert analog audio input into digital sound. The audio module 170 can also be used for sound encoding and decoding. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.
[0060] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or hands-free calls through the speaker 170A.
[0061] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the voice can be listened to by placing the receiver 170B close to the human ear.
[0062] The microphone 170C, also known as a "microphone" or "transmitter", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak close to the microphone 170C with their mouth to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In some other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.
[0063] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0064] The motor 191 can generate a vibration prompt. The motor 191 can be used for an incoming call vibration prompt or a touch vibration feedback. For example, touch operations applied to different applications (such as taking pictures, playing audio, etc.) can correspond to different vibration feedback effects. For touch operations applied to different areas of the display screen 194, the motor 191 can also correspond to different vibration feedback effects. Different application scenarios (such as time reminder, receiving information, alarm clock, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0065] The indicator 192 can be an indicator light, which can be used to indicate the charging state, power change, or can also be used to indicate messages, missed calls, notifications, etc.
[0066] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present invention, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the electronic device 100.
[0067] Figure 2 It is a software structure block diagram of the electronic device 100 in the embodiments of the present invention.
[0068] The layered architecture divides software into several layers, and each layer has clear roles and divisions of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0069] The application layer may include a series of application packages.
[0070] As Figure 2 shown, the application packages may include applications such as the camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.
[0071] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.
[0072] As Figure 2 shown, the application framework layer may include the window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0073] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0074] The content provider is used to store and obtain data, and make this data accessible to applications. The data may include videos, images, audio, dialed and answered calls, browsing history and bookmarks, phone book, etc.
[0075] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a short message notification icon may include a view for displaying text and a view for displaying pictures.
[0076] The phone manager is used to provide the communication function of the electronic device 100. For example, the management of call states (including connection, disconnection, etc.).
[0077] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, etc.
[0078] The notification manager enables an application to display notification information in the status bar. It can be used to convey messages of the notification type, and can disappear automatically after a short stay without user interaction. For example, the notification manager is used to inform that a download is completed, a message is reminded, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or a scroll bar text, such as the notification of a background running application, or a notification that appears on the screen in the form of a dialogue window. For example, it can prompt text information in the status bar, emit a prompt sound, vibrate the electronic device, blink the indicator light, etc.
[0079] Android Runtime includes the core libraries and the virtual machine. Android runtime is responsible for the scheduling and management of the Android system.
[0080] The core libraries include two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core libraries of Android.
[0081] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of the object life cycle, stack management, thread management, security and exception management, and garbage collection.
[0082] The system libraries can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (such as: OpenGL ES), 2D graphics engine (such as: SGL), etc.
[0083] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.
[0084] The media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0085] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.
[0086] The 2D graphics engine is the drawing engine for 2D drawing.
[0087] The kernel layer is the layer between the hardware and the software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.
[0088] For the sake of easy understanding, the following embodiments of this application will be based on having Figure 1 and Figure 2Taking the electronic device with the structure shown as an example, in combination with the accompanying drawings and application scenarios, the method provided in the embodiments of the present application will be specifically described. It should be noted that although the software structure of the electronic device is based on Figure 2 as an example, Figure 2 the software structure shown is only a schematic example, and the software structures of other operating systems are also applicable to the wake-up method provided in the embodiments of the present application.
[0089] For the convenience of description, taking the electronic device as a smart speaker and the smart speaker being in a home environment as an example, the wake-up method provided in the embodiments of the present application will be clarified. Figure 3 This is a schematic diagram of a scenario of the wake-up method provided in the embodiments of the present application. As Figure 3 shown, in addition to the smart speaker, the home environment is also provided with other devices with functions of playing external sounds, such as a TV and a traditional speaker, and furniture such as a sofa and a dining table. Users can move around near the sofa, dining table, etc., and wake up the smart speaker by uttering a voice containing a wake-up word. The smart speaker can also be placed in other scenarios. For example, shopping malls, office environments, etc. By executing the wake-up method provided in the embodiments of the present application, the probability of false wake-up of the smart speaker can also be reduced. Hereinafter, the specific implementation of the wake-up method provided in the embodiments of the present application will be described.
[0090] When the smart speaker is not awakened and in the standby state, it picks up the sound in the environment to obtain the sound. This sound includes the sound of the target speaker, that is, the user, and also includes the noise signal in the environment. Therefore, generally, the received sound is denoised to obtain a clean sound, which is used as the sound for triggering the execution of the wake-up method provided in the embodiments of the present application.
[0091] In the wake-up method provided in the embodiments of the present application, at least one of a wake-up azimuth set, a false wake-up azimuth set, and a voiceprint set is set in the smart speaker. Among them, the false wake-up azimuth set includes: the false wake-up azimuth and the confidence level of the false wake-up azimuth. Hereinafter, an element in the false wake-up azimuth set is represented by (false wake-up azimuth, confidence level). The false wake-up azimuth is used to record the sound source azimuth of the sound that does not wake up the smart speaker. The confidence level of the false wake-up azimuth is used to describe the probability of uttering a voice to wake up the smart speaker at the false wake-up azimuth. The confidence level can be identified by the magnitude of a numerical value. For example, the larger the numerical value, the higher the probability, and the smaller the numerical value, the lower the probability. The azimuth in the embodiments of the present application refers to the direction and position relative to the smart speaker. For example, the sound source azimuth refers to the direction and position of the sound source relative to the smart speaker.
[0092] The wake-up orientation set includes: the wake-up orientation and the confidence level of the wake-up orientation. Hereinafter, an element in the wake-up orientation set is represented by (wake-up orientation, confidence level). The wake-up orientation is used to record the sound source orientation of the sound that wakes up the smart speaker. The confidence level of the wake-up orientation is used to describe the probability that a voice for waking up the smart speaker is emitted at the wake-up orientation.
[0093] The voiceprint set includes: the wake-up voiceprint and the confidence level of the wake-up voiceprint. The confidence level of the wake-up voiceprint can be represented by the number of hits of the wake-up voiceprint. Hereinafter, an element in the voiceprint set is represented by (wake-up voiceprint, confidence level). The wake-up voiceprint is used to record the user's voiceprint of the sound that wakes up the smart speaker. The confidence level of the wake-up voiceprint is used to record the probability that the sound with this wake-up voiceprint wakes up the smart speaker. The number of hits is used to record the number of times the sound with this wake-up voiceprint wakes up the smart speaker. The user's voiceprint and the wake-up voiceprint can be represented by the parameter values of the voiceprint feature parameters. The above-mentioned voiceprint feature parameters can include, for example, but are not limited to, intensity, wavelength, frequency, rhythm, etc. The parameter values of at least one voiceprint feature parameter are different between different voiceprints.
[0094] Possible representation methods and calculation methods for the sound source orientation are described as follows: Optionally, a coordinate system of the smart speaker can be established. For example, the origin of the coordinate system can be the physical center point of the smart speaker, and the positive direction of the x-axis can be the direction horizontally pointing directly in front of the smart speaker. The method of establishing this coordinate system is only an example and is not used to limit the method of establishing the coordinate system of the smart speaker. The sound source orientation can be identified by distance and angle in the above coordinate system. Specifically, the distance of the sound source orientation can be used to record: the distance between the sound source of the sound and the origin of the coordinate system of the smart speaker. The angle can be used to record: the included angle between the ray pointing from the origin of the coordinate system of the smart speaker to the sound source of the sound and the positive direction of the x-axis of the smart speaker. Optionally, a parameter of the height dimension can be further added to the sound source orientation. The height can be used to record: the vertical distance between the sound source of the sound and the origin of the coordinate system. Information such as the distance and angle of the sound source orientation can be calculated by the smart speaker based on relevant sound source localization methods. The sound source localization method can calculate the relative position between the sound source and the smart speaker based on a microphone array composed of at least 2 microphones set in the smart speaker. Such as distance and angle. Specifically, the sound source localization method can include but is not limited to: the controllable beamforming technology based on the maximum output power, the high-resolution spectrogram estimation technology, and the sound source localization technology based on the time-delay estimation (TDE) of sound, etc. Taking the algorithm based on TDE as an example, its core lies in the accurate estimation of the propagation delay, and the propagation delay is generally obtained by performing cross-correlation processing on the sounds picked up by the microphone array of the smart speaker. After that, the distance between the smart speaker and the sound source can be calculated by methods such as simple delay summation, geometric calculation, or directly using the cross-correlation result for controllable power response search. The specific algorithms are not elaborated one by one in the embodiments of this application.
[0095] Initial setting of the set: In the initial situation, for example, when the smart speaker leaves the factory unused or is restored to the factory settings, at least one of the preset false wake-up orientation set, wake-up orientation set, and voiceprint set in the smart speaker can be empty. During the user's use of the smart speaker, the user can set at least one of the false wake-up orientation set, wake-up orientation set, and voiceprint set based on the environment where the smart speaker is located, or can also not set it. If the user does not set it, it can reduce the user operation and improve the user experience.
[0096] An example of the method for setting the above set is given: Since the false wake-up orientation records the sound source orientation of the sound that does not wake up the smart speaker, the false wake-up orientation generally corresponds to the orientation of other devices that can emit sound in the environment relative to the smart speaker. Based on this, the false wake-up orientation can be set based on the orientation of other devices that can emit sound in the environment relative to the smart speaker, and the user or the smart speaker can set an initial confidence level for this false wake-up orientation. Figure 3Taking the home environment shown as an example, when the smart speaker has a display screen, a setting interface for the mis-awakening direction can be provided for the user on the display screen of the smart speaker. For example Figure 4 As shown, the user can set a mis-awakening direction and the confidence level of this mis-awakening direction based on the relative position between the TV and the smart speaker; set a mis-awakening direction and the confidence level of this mis-awakening direction based on the relative position between the traditional speaker and the smart speaker; then click the "OK" control; correspondingly, the smart speaker detects the user's operation on the "OK" control in the setting interface, obtains information such as the mis-awakening direction and the confidence level in the setting interface, and saves them in the mis-awakening direction set. Optionally, if the smart speaker does not include a display screen or the display screen is not convenient for the user to operate, this setting interface can be displayed to the user by other devices associated with the smart speaker (such as the user's smartphone, etc.), and the other devices will send the mis-awakening direction, the confidence level and other information obtained from the setting interface to the smart speaker.
[0097] Since the wake-up direction records the sound source direction of the sound that wakes up the smart speaker, the wake-up direction generally corresponds to the direction of the position in the environment where the user often emits the voice to wake up the smart speaker relative to the smart speaker. Based on this, the user can set the wake-up direction based on the direction of the position in the environment where the user often wakes up the smart speaker relative to the smart speaker, and the user or the smart speaker sets an initial confidence level for this wake-up direction. Taking Figure 3 the home environment shown as an example, generally the user often moves around near the sofa, dining table, etc. and wakes up the smart speaker. Therefore, one or more wake-up directions and the corresponding confidence levels can be set based on the direction of the position on the sofa relative to the smart speaker, and one or more wake-up directions and the corresponding confidence levels can be set based on the direction of the position near the dining table, such as the position of the dining chair, relative to the smart speaker. The specific setting method can refer to Figure 4 the setting method of the mis-awakening direction shown, which will not be elaborated here.
[0098] The wake-up voiceprint can be set by the user by recording voice. Correspondingly, the smart speaker obtains the user's voiceprint according to the sound obtained from the recorded voice and sets it as the wake-up voiceprint. The user or the smart speaker sets the initial confidence level of this wake-up voiceprint. For example, if the confidence level is the number of hits, the initial confidence level can be 0.
[0099] Update of the set: During the use of the smart speaker, when the smart speaker is awakened, the wake-up direction set can be updated according to the sound source direction of the sound that wakes up the smart speaker; the voiceprint set can be updated according to the user's voiceprint extracted from this sound; when the smart speaker is not awakened, the mis-awakening direction set can be updated according to the sound source direction of the sound that does not wake up the smart speaker.
[0100] Based on the sound that wakes up the smart speaker, calculate the sound source direction of the sound, and determine whether the set of wake-up directions includes the sound source direction of the sound. If it includes, increase the confidence level of the wake-up direction corresponding to the sound source direction; if it does not include, add the sound source direction as a wake-up direction to the set of wake-up directions, and set an initial confidence level for the newly added wake-up direction. Among them, when the intelligent voice signal determines whether the set of wake-up directions includes the sound source direction of the sound, the sound source direction can be exactly the same as a certain wake-up direction, or there can be a certain deviation. For example, when the wake-up direction and the sound source direction are respectively represented by (distance, angle), a distance threshold and an angle threshold can be preset respectively. If the distance difference between the sound source direction and the wake-up direction 1 satisfies the distance threshold and the angle difference satisfies the angle threshold, it can be determined that the set of wake-up directions includes the sound source direction. The wake-up direction 1 can be called the wake-up direction corresponding to the sound source direction, or the wake-up direction including the sound source direction. Correspondingly, increase the confidence level of the wake-up direction 1 corresponding to the sound source direction. It should be noted that the embodiments of the present application do not limit the set value of the initial confidence level. The embodiments of the present application also do not limit the amplitude of the increase in the confidence level each time the confidence level of the wake-up direction is increased. For example, the amplitude can be a certain fixed value, or a fixed percentage of the confidence level, etc. Similarly, the embodiments of the present application do not limit the specific values of the preset distance threshold and angle threshold. The distance threshold and the angle threshold can be determined based on the accuracy of the wake-up method and the accuracy of the sound source direction calculation method, etc. Specifically, the higher the accuracy of the wake-up method, the smaller the values of the distance threshold and the angle threshold generally are; the higher the accuracy of the sound source direction calculation method, the smaller the values of the distance threshold and the angle threshold generally are. In addition, the setting of the distance threshold and the angle threshold can expand the wake-up direction in the set of wake-up directions from a point to a region, and the distance threshold and the angle threshold can be set based on the size of the desired expanded region. The distance threshold and the angle threshold can be adjusted by the user of the smart speaker according to needs.
[0101] Based on the sound that wakes up the smart speaker, the smart speaker can extract the user's voiceprint of the sound, and determine whether the wake-up voiceprint in the set of voiceprints includes the extracted user's voiceprint. If it includes, increase the confidence level of the wake-up voiceprint; otherwise, add the user's voiceprint as a wake-up voiceprint to the set of voiceprints, and set a confidence level for the newly added wake-up voiceprint. Similar to the judgment of the set of wake-up directions, a certain error can also be allowed between the user's voiceprint and the wake-up voiceprint when determining whether the set of voiceprints includes the extracted user's voiceprint. For example, a threshold can be set for each voiceprint feature included in the voiceprint. As long as the difference between the value of each voiceprint feature of the user's voiceprint and the value of the corresponding voiceprint feature of a certain wake-up voiceprint is less than the threshold corresponding to the voiceprint feature, it can be considered that the set of voiceprints includes the user's voiceprint, and the above-mentioned certain wake-up voiceprint is the wake-up voiceprint corresponding to the user's voiceprint.
[0102] Based on the sound of an unawakened smart speaker, the intelligent voice signal calculates the sound source direction of the sound. The intelligent voice signal determines whether the false awakening direction set includes the sound source direction of the sound. If it includes, the confidence level of the false awakening direction corresponding to the sound source direction is reduced; if it does not include, the sound source direction is added to the false awakening direction set as a false awakening direction, and an initial confidence level is set for the newly added false awakening direction. The implementation of the update of the false awakening direction set can refer to the relevant description in the update of the awakening direction set, which will not be elaborated here.
[0103] The awakening method in the embodiments of the present application determines whether to execute the awakening process based on the awakening word confidence level output by the awakening word model, the false awakening direction set and / or the awakening direction set, and the voiceprint set, so as to reduce the false awakening probability. The following specifically describes the awakening method.
[0104] In one implementation, the smart speaker includes a pickup and a speaker. Among them, the pickup includes a microphone array, and the microphone array includes multiple microphones.
[0105] As Figure 5 shown, the smart speaker is preset with an awakening direction set (also referred to as the first direction set), a false awakening direction set (also referred to as the second direction set), and a voiceprint set (also referred to as the first voiceprint set). The awakening method in the embodiments of the present application may include:
[0106] Step 501: The smart speaker picks up the sound in the environment to obtain the sound.
[0107] Since the smart speaker generally picks up the sound in the environment continuously, the smart speaker generally divides the continuously picked-up sound into audio segments according to a certain duration. The sound in the embodiments of the present application generally refers to the divided audio segments. The specific duration of the audio segment is not limited in the embodiments of the present application.
[0108] In order to reduce the influence of noise on subsequent processing, before executing step 502, the smart speaker generally performs noise reduction processing on the sound to suppress the noise signal in the sound and obtain a relatively clean sound. In this way, the sound used in step 502 is generally the sound after noise reduction processing.
[0109] Since the smart speaker picks up sound continuously, in order to reduce the data processing amount and power consumption of the smart speaker, preset conditions such as a sound intensity threshold can be set for the sound picked up. Only the sound that meets the preset conditions will calculate the awakening word confidence level based on the awakening word model to trigger subsequent processing. The specific preset conditions are not limited in the embodiments of the present application.
[0110] Step 502: The smart speaker calculates the awakening word confidence level of the sound based on the awakening word model.
[0111] The wake word confidence is used to describe the probability that the voice includes the wake word voice.
[0112] Step 503: The smart speaker determines whether the wake word confidence is less than the first threshold; if the wake word confidence is not less than the first threshold, step 504 is executed.
[0113] Further, step 503 further includes: if it is less than the first threshold, the wake-up process is not executed, and the false wake-up azimuth set is updated according to the sound source azimuth of the voice, and this branch process ends.
[0114] It should be noted that the judgment in step 503 can also be executed by the wake word model, so that the wake word model can output two parameters, namely the judgment result of whether to wake up and the wake word confidence. The embodiments of the present application do not make limitations.
[0115] The wake word confidence is used to describe the probability that the voice includes the wake word voice. The higher the wake word confidence, the greater the probability that the voice includes the wake word voice. The wake-up method of the embodiments of the present application further executes the following steps 504 to step 511 to further determine whether to execute the wake-up process, so as to implement the screening of false wake-ups and reduce the false wake-up probability.
[0116] Step 504: The smart speaker determines whether the wake word confidence is less than the second threshold, and the second threshold is greater than the first threshold. If it is not less than the second threshold, the wake-up process is executed, and step 511 is executed; if it is less than the second threshold, step 505 is executed.
[0117] In the embodiments of the present application, the situation where the wake word confidence is not less than the first threshold is further divided into two types by the second threshold: if the wake word confidence is not less than the second threshold, it means that the probability that the voice includes the wake word voice is high and the probability of false wake-up is low, so the wake-up process is directly executed to wake up the smart speaker; if the wake word confidence is less than the second threshold and not less than the first threshold, it means that the probability that the voice includes the wake word voice is relatively low and the probability of false wake-up is relatively high, so the following steps 506 to step 509 are executed to further determine whether to execute the wake-up process in combination with the wake-up azimuth set, the false wake-up azimuth set, or the voiceprint set. For example, the value range of the wake word confidence is (0, 100), the first threshold is 30, and the second threshold is 80. Correspondingly, if the wake word confidence is less than 30, the wake-up process is not executed; if the wake word confidence is not less than 80, the wake-up process is directly executed; if the wake word confidence is less than 80 and not less than 30, the following steps 505 to step 509 are executed to further screen out possible false wake-ups.
[0118] The wake-up orientation set, the false wake-up orientation set, or the voiceprint set may include at least one set element, and each set element includes at least two units. For example, the set element included in the wake-up orientation set includes the wake-up orientation and the confidence level corresponding to the wake-up orientation; the set element included in the false wake-up orientation set includes the false wake-up orientation and the confidence level corresponding to the false wake-up orientation; the set element included in the voiceprint set includes the voiceprint and the confidence level corresponding to the voiceprint.
[0119] It should be noted that there is no execution order restriction between the steps of executing the wake-up process in step 504 and step 511. Figure 5 Taking the example of executing the wake-up process first and then executing step 511.
[0120] Both the first threshold and the second threshold can be preset.
[0121] Step 505: The smart speaker calculates the sound source orientation of the sound.
[0122] The calculation method of the sound source orientation has been described above and will not be elaborated here.
[0123] Step 506: The smart speaker determines whether the false wake-up orientation set includes the sound source orientation of the sound, and determines whether the wake-up orientation set includes the sound source orientation of the sound; if only the false wake-up orientation set includes the sound source orientation, execute step 507; if only the wake-up orientation set includes the sound source orientation, execute step 508; if it does not belong to the above two situations, execute step 509.
[0124] Step 507: The smart speaker determines whether to execute the wake-up process according to the confidence level of the false wake-up orientation corresponding to the sound source orientation; if yes, execute the wake-up process and execute step 511; if no, do not execute the wake-up process, update the false wake-up orientation set according to the sound source orientation of the sound, and this branch process ends.
[0125] Among them, the smart speaker determines whether to execute the wake-up process according to the confidence level of the false wake-up orientation corresponding to the sound source orientation, which may include:
[0126] If the confidence level of the false wake-up orientation is less than threshold a, it is determined not to execute the wake-up process;
[0127] If the confidence level of the false wake-up orientation is not less than threshold a, it is determined to execute the wake-up process;
[0128] If the confidence level of the false wake-up orientation is less than threshold a, it means that the probability that the sound is a noise signal is relatively high, so it is determined not to execute the wake-up process, that is, not to wake up the smart speaker, thereby reducing the false wake-up probability.
[0129] Step 508: The smart speaker determines whether to execute the wake-up process based on the confidence level of the wake-up direction corresponding to the sound source direction; if yes, execute the wake-up process and perform Step 511; if no, do not execute the wake-up process, update the false wake-up direction set according to the sound source direction of the sound, and this branch process ends.
[0130] Among them, the smart speaker determines whether to execute the wake-up process based on the confidence level of the wake-up direction corresponding to the sound source direction, which may include:
[0131] If the confidence level of the wake-up direction is less than threshold b, it is determined not to execute the wake-up process;
[0132] If the confidence level of the wake-up direction is not less than threshold b, it is determined to execute the wake-up process;
[0133] If the confidence level of the wake-up direction is less than threshold b, it indicates that the probability that the sound is a noise signal is relatively high, so it is determined not to execute the wake-up process, that is, not to wake up the smart speaker, thereby reducing the false wake-up probability.
[0134] Step 509: The smart speaker extracts the user voiceprint of the sound and determines whether the wake-up voiceprint in the voiceprint set includes the extracted user voiceprint; if it includes, perform Step 510; if it does not include, do not execute the wake-up process; update the false wake-up direction set according to the sound source direction of the sound, and this branch process ends;
[0135] Step 510: The smart speaker determines whether to execute the wake-up process based on the confidence level of the wake-up voiceprint corresponding to the user voiceprint; if yes, execute the wake-up process and perform Step 511; if no, do not execute the wake-up process, update the false wake-up direction set according to the sound source direction of the sound, and this branch process ends.
[0136] Among them, after the judgment result is yes, the execution order between the wake-up process and Step 511 is not restricted.
[0137] Among them, the smart speaker determines whether to execute the wake-up process based on the confidence level of the wake-up voiceprint corresponding to the user voiceprint, which may include:
[0138] If the confidence level of the wake-up voiceprint is less than threshold c, it is determined not to execute the wake-up process;
[0139] If the confidence level of the wake-up voiceprint is not less than threshold c, it is determined to execute the wake-up process;
[0140] If the confidence level of the wake-up voiceprint is less than threshold c, it indicates that the possibility that the sound is made by a person who does not often appear in the environment is relatively high, so it is determined not to execute the wake-up process, that is, not to wake up the smart speaker, thereby reducing the false wake-up probability.
[0141] Step 511: The smart speaker updates the wake-up orientation set according to the sound source orientation of the sound and updates the voiceprint set according to the user's voiceprint of the sound, and this branch process ends.
[0142] The implementation of this step can refer to the foregoing description of set update, which will not be elaborated here.
[0143] It should be noted that Figure 5 In the illustrated embodiment, the second threshold may not be set, that is, step 504 is not executed, and step 505 is directly executed. Or, in another possible implementation, the judgment of step 504 may be moved to Figure 5 Steps 507, 508, and 510 in are executed, and the smart speaker combines the wake-up word confidence to determine whether to execute the wake-up process. The specific judgment criteria can refer to Figure 5 The judgment criteria shown. Taking step 507 as an example, step 507 will be replaced with: The smart speaker determines whether to execute the wake-up process according to the wake-up word confidence and the confidence of the mis-wake-up orientation corresponding to the sound source orientation; if the judgment result is yes, the wake-up process is executed, and step 511 is executed; if the judgment result is no, the wake-up process is not executed, and step 512 is executed. At this time, when the smart speaker in step 507 determines whether to execute the wake-up process according to the wake-up word confidence and the confidence of the mis-wake-up orientation corresponding to the sound source orientation, it may include:
[0144] If the wake-up word confidence is not less than the second threshold, it is judged to execute the wake-up process;
[0145] If the wake-up word confidence is less than the second threshold and the confidence of the mis-wake-up orientation is less than the first orientation confidence threshold, it is judged not to execute the wake-up process;
[0146] If the wake-up word confidence is less than the second threshold and the confidence of the mis-wake-up orientation is not less than the first orientation confidence threshold, it is judged to execute the wake-up process.
[0147] It should be noted that Figure 5 In, it is taken as an example that each time the smart speaker determines to execute the wake-up process, the wake-up orientation set and the voiceprint set are updated, and each time it determines not to execute the wake-up process, the mis-wake-up orientation set is updated. However, considering reducing the data processing amount and power consumption of the smart speaker, etc., it is not necessary to update the above sets every time the smart speaker determines to execute or not execute the wake-up process, but to update the above sets after selecting some times of judgment based on a certain rule. The embodiments of the present application do not make limitations.
[0148] After the smart speaker determines to execute the wake-up process, it updates the wake-up azimuth set and the voiceprint set; after determining not to execute the wake-up process, it updates the false wake-up azimuth set; in this way, during the gradual use of the smart speaker, the wake-up azimuths recorded in the wake-up azimuth set can correspond to the positions where the user often emits wake-up voice commands in the environment, the false wake-up azimuths recorded in the false wake-up azimuth set can correspond to the positions of other devices that can emit sounds in the environment, and the highly confident wake-up voiceprints recorded in the voiceprint set correspond to the user voiceprints that often wake up the smart speaker, so that the wake-up method of the embodiments of the present application can better achieve the effect of reducing the false wake-up probability.
[0149] For example: Suppose the smart speaker is placed in a home environment for use after leaving the factory. The smart speaker is preset with a wake-up azimuth set, a false wake-up azimuth set, and a voiceprint set, and the three sets are all empty; let the first threshold be 0.4, the second threshold be 0.7, the threshold a be 0.5, the threshold b be 0.6, and the threshold c be 5; then,
[0150] Suppose the TV placed in the environment makes a relatively loud sound. The smart speaker picks up the sound 1 of the TV. Based on the wake-up word model, the wake-up word confidence of the sound 1 is calculated to be 0.1, which is less than the preset first threshold of 0.4. Then, execute the branch where the judgment result in step 503 is yes, do not execute the wake-up, add the sound source azimuth 1 of the sound 1 to the false wake-up azimuth set, obtain the false wake-up azimuth 1, and set its initial confidence, for example, 0.8; since the false wake-up probability is generally very small, therefore, the smart speaker picks up the sound 2 of the TV again, and then calculates that the wake-up word confidence of the sound 2 is 0.2, and executes the branch where the judgment result in 503 is yes, and reduces the confidence of the false wake-up azimuth 1 in the false wake-up azimuth set; as the smart speaker executes the above process multiple times, the confidence of the false wake-up azimuth 1 is reduced. Once this confidence is reduced below the threshold a of 0.5, for example, 0.45, then even if occasionally the smart speaker picks up the sound n emitted by the TV, and calculates that the wake-up word confidence of the sound n is a value between the first threshold of 0.4 and the second threshold of 0.7, such as 0.55, the smart speaker sequentially executes step 503, step 504, and step 505. According to the confidence of the false wake-up azimuth 1 of 0.48, it is judged that it is less than the threshold a of 0.5, and the wake-up process will not be executed, so as to screen out possible false wake-up situations from the situation of executing the wake-up process in the prior art, thereby reducing the false wake-up probability;
[0151] Even if a false wake-up caused by the TV sound has occurred before the confidence level of the false wake-up direction 1 drops below the threshold a of 0.5, for example, the smart speaker picks up the TV sound to obtain sound m, and the calculated wake-up word confidence level is 0.65. Steps 503 to 505 are sequentially executed. According to the confidence level of the false wake-up direction 1, such as 0.55, it is judged that it is not less than the threshold a of 0.5, and the wake-up process is executed. At this time, the wake-up direction set and the voiceprint set will be updated, so that the wake-up direction set also includes the sound source direction 1. However, during the above use process, the smart speaker will also update the wake-up direction set and the voiceprint set according to the sound when the user wakes up the smart speaker. Among the wake-up voiceprints recorded in the voiceprint set, the confidence level of the user voiceprint that often wakes up the smart speaker gradually increases. When the smart speaker subsequently obtains a sound with a wake-up word confidence level between 0.4 and 0.7 at the sound source direction where the TV is located, it can further judge whether to execute the wake-up process according to this voiceprint set, so as to screen out possible false wake-ups and reduce the false wake-up probability.
[0152] Figure 5 In the wake-up method of the embodiment of the present application shown, after calculating the wake-up word confidence level of the sound based on the wake-up word model, it is further combined with the wake-up direction set, the false wake-up direction set and the voiceprint set to judge whether to wake up the smart speaker, so as to further screen out possible false wake-ups when the judgment result output by the wake-up word model is to wake up the smart speaker, thereby reducing the false wake-up probability of the smart speaker and improving the user experience.
[0153] Alternatively, the wake-up direction set, the false wake-up direction set and the voiceprint set may also be non-preset, but created with machine learning; after creation, continue to enrich and adjust according to machine learning.
[0154] Different from Figure 5 taking the preset wake-up direction set, false wake-up direction set and voiceprint set in the smart speaker as an example, in Figure 6 the embodiment shown, taking the preset false wake-up direction set and voiceprint set in the smart speaker as an example, at this time, steps 506 to 511 are replaced by the following steps 601 to 605, specifically:
[0155] Step 601: The smart speaker judges whether the false wake-up direction set includes the sound source direction of the sound; if it includes, execute step 602; if it does not include, execute step 603;
[0156] The implementation of this step can refer to the relevant judgment method in the update of the false wake-up direction set above, and will not be elaborated here.
[0157] Step 602: The smart speaker determines whether to execute the wake-up process based on the confidence level of the false wake-up direction corresponding to the sound source direction; if the determination result is yes, execute the wake-up process and perform Step 605; if the determination result is no, do not execute the wake-up process, update the false wake-up direction set according to the sound source direction of the sound, and end this branch process.
[0158] For the implementation of this step, please refer to the description in Step 507 and will not be elaborated here.
[0159] Step 603: The smart speaker extracts the user voiceprint of the sound and determines whether the wake-up voiceprints in the voiceprint set include the extracted user voiceprint; if it includes, perform Step 604; if it does not include, do not execute the wake-up process, update the false wake-up direction set according to the sound source direction of the sound, and end this branch process.
[0160] Step 604: The smart speaker determines whether to execute the wake-up process based on the confidence level of the wake-up voiceprint corresponding to the user voiceprint; if the determination result is yes, execute the wake-up process and perform Step 605; if the determination result is no, do not execute the wake-up process, update the false wake-up direction set according to the sound source direction of the sound, and end this branch process.
[0161] Step 605: The smart speaker updates the voiceprint set according to the user voiceprint of the sound.
[0162] For the implementation of this step, please refer to the description in Step 511 and will not be elaborated here.
[0163] Figure 6 In the wake-up method of the embodiment of the present application shown, after calculating the wake-up word confidence level of the sound based on the wake-up word model, further combine the false wake-up direction set and the voiceprint set to determine whether to wake up the smart speaker, so as to further screen out possible false wake-ups in the case where the judgment result output by the wake-up word model is to wake up the smart speaker, thereby reducing the false wake-up probability of the smart speaker and improving the user experience.
[0164] Different from Figure 6 In the wake-up method shown, the preset false wake-up direction set and voiceprint set in the smart speaker are taken as examples, Figure 7 In the wake-up method shown, the preset wake-up direction set and voiceprint set in the smart speaker are taken as examples; different from Figure 6 The main difference is that the false wake-up direction set is replaced by the wake-up direction set, and the step of updating the false wake-up direction set is omitted. In Step 705, the wake-up direction set is updated according to the sound source direction of the sound, and the voiceprint set is updated according to the user voiceprint of the sound.
[0165] For the implementation of determining whether to execute the wake-up process based on the confidence level of the wake-up direction in Step 702, please refer to the description in Step 508 and will not be elaborated here.
[0166] When step 705 is executed for the first time, update the first azimuth set and the first voiceprint set, including: creating the first azimuth set and the first voiceprint set; incorporating the azimuth that wakes up the electronic device into the first azimuth set and assigning an initial first azimuth confidence level to the incorporated azimuth; and incorporating the voiceprint that wakes up the electronic device into the first voiceprint set and assigning an initial first voiceprint confidence level to the incorporated voiceprint.
[0167] When step 705 is executed subsequently, update the first azimuth set and the first voiceprint set, including at least one of the following: creating a new first azimuth in the first azimuth set and assigning an initial first azimuth confidence level to the newly created first azimuth; creating a new first voiceprint in the first voiceprint set and assigning an initial first voiceprint confidence level to the newly created first voiceprint; increasing the first azimuth confidence level corresponding to an existing first azimuth that is matched in the first azimuth set; increasing the first voiceprint confidence level corresponding to an existing first voiceprint that is matched in the first voiceprint set.
[0168] Although taking step 705 in Figure 7 as an example for illustration, Figure 6 the updating of the mis-awakening azimuth set, the updating of the voiceprint set, etc. in steps 602 to 605 in Figure 5 is similar thereto, and the updating of the mis-awakening azimuth set, the updating of the voiceprint set, etc. in steps 507 to 511 in
[0169] Figure 7 is also similar thereto; they will not be elaborated one by one here.
[0169] Figure 7 The wake-up method of the embodiment of the present application shown determines whether to wake up the smart speaker by further combining the wake-up azimuth set and the voiceprint set based on the wake-up word confidence level calculated for the sound. Thus, when the judgment result output by the wake-up word model is to wake up the smart speaker, possible mis-awakenings are further screened out, thereby reducing the mis-awakening probability of the smart speaker and improving the user experience.
[0170] Figure 8 is a schematic flowchart of another embodiment in the wake-up method provided by the embodiment of the present application. This method can be applied to an electronic device such as the above-mentioned smart speaker. This method may include:
[0171] Step 801: Receive a sound and calculate the wake-up word confidence level of the sound; the wake-up word confidence level is used to describe the probability that the sound includes the wake-up word sound.
[0172] Step 802: If the wake-up word confidence level is greater than or equal to the first threshold, calculate the sound source azimuth of the sound.
[0173] Step 803: Determine whether the sound source direction is in the first direction set or the second direction set; wherein, the first direction set includes several first directions, and the first directions are used to record the sound source directions of the sounds that do not wake up the smart speaker; the second direction set includes several second directions, and the second directions are used to record the sound source directions of the sounds that wake up the smart speaker.
[0174] Step 804: If the sound source direction is only in the first direction set, determine whether to wake up the smart speaker according to the confidence level of the first direction corresponding to the sound source direction, and the confidence level of the first direction is used to describe the probability of the voice for waking up the smart speaker being emitted at the first direction.
[0175] Step 805: If the sound source direction is only in the second direction set, determine whether to wake up the smart speaker according to the confidence level of the second direction corresponding to the sound source direction, and the confidence level of the second direction is used to describe the probability of the voice for waking up the smart speaker being emitted at the second direction.
[0176] Wherein, the confidence level of the wake-up word can correspond to the above-mentioned confidence level of the wake-up word, the first direction can correspond to the above-mentioned mis-wake-up direction, and the second direction can correspond to the above-mentioned wake-up direction.
[0177] In a possible implementation manner, it may further include:
[0178] If the sound source direction is in both the first direction set and the second direction set, or the sound source direction is not in the first direction set and the second direction set, extract the user's voiceprint from the sound;
[0179] Determine whether the user's voiceprint is included in the first voiceprint set; the first voiceprint set includes first voiceprints, and the first voiceprints are used to record the user's voiceprints of the sounds that wake up the electronic device;
[0180] Determine whether to wake up the electronic device according to the confidence level of the first voiceprint corresponding to the user's voiceprint.
[0181] In a possible implementation manner, before calculating the sound source direction of the sound, it may further include: determining that the confidence level of the wake-up word is less than a second threshold; the second threshold is greater than the first threshold.
[0182] In a possible implementation manner, determining whether to wake up the electronic device according to the confidence level of the first direction corresponding to the sound source direction may include:
[0183] Determine whether the confidence level of the first direction corresponding to the sound source direction is less than threshold a;
[0184] If it is less than threshold a, the determination result is not to wake up the electronic device;
[0185] If it is not less than threshold a, the determination result is to wake up the electronic device.
[0186] In a possible implementation, determining whether to wake up the electronic device based on the confidence level of the second azimuth corresponding to the sound source azimuth may include:
[0187] Determining whether the confidence level of the second azimuth corresponding to the sound source azimuth is less than threshold b;
[0188] If it is less than threshold b, the determination result is not to wake up the electronic device;
[0189] If it is not less than threshold b, the determination result is to wake up the electronic device.
[0190] In a possible implementation, determining whether to wake up the electronic device based on the confidence level of the first voiceprint corresponding to the user's voiceprint may include:
[0191] Determining whether the confidence level of the first voiceprint corresponding to the user's voiceprint is less than threshold c;
[0192] If it is less than threshold c, the determination result is not to wake up the electronic device;
[0193] If it is not less than threshold c, the determination result is to wake up the electronic device.
[0194] In a possible implementation, it may further include:
[0195] If the determination result is to wake up the electronic device and the sound source azimuth of the sound is included in the second azimuth set, increase the confidence level of the second azimuth corresponding to the sound source azimuth;
[0196] If the determination result is to wake up the electronic device and the sound source azimuth of the sound is not included in the second azimuth set, store the sound source azimuth as the second azimuth in the second azimuth set and set an initial confidence level for this second azimuth.
[0197] In a possible implementation, it may further include:
[0198] If the determination result is to wake up the electronic device and the user's voiceprint of the sound is included in the first voiceprint set, increase the confidence level of the first voiceprint corresponding to the user's voiceprint;
[0199] If the determination result is to wake up the electronic device and the user's voiceprint of the sound is not included in the first voiceprint set, store the user's voiceprint as the first voiceprint in the first voiceprint set and set an initial confidence level for this first voiceprint.
[0200] In a possible implementation, it may further include:
[0201] If the determination result is not to wake up the electronic device and the sound source azimuth of the sound is included in the first azimuth set, decrease the confidence level of the first azimuth including the sound source azimuth;
[0202] If the judgment result is not to wake up the electronic device, and the sound source direction of the sound is not included in the first direction set, store the sound source direction as the first direction in the first direction set, and set an initial confidence level for the first direction.
[0203] Figure 8 For the specific implementation, reference can be made to Figure 5 the embodiments shown, which will not be elaborated here.
[0204] Figure 9 FIG. is a flowchart of another embodiment of the wake-up method of the present application. This method can be applied to an electronic device such as the above-mentioned smart speaker. This method may include:
[0205] Step 901: Receive a sound and calculate the wake-up word confidence level of the sound; the wake-up word confidence level is used to describe the probability that the sound includes the wake-up word sound;
[0206] Step 902: If the wake-up word confidence level is greater than or equal to the first threshold, calculate the sound source direction of the sound;
[0207] Step 903: Determine whether the sound source direction is in the first direction set; wherein, the first direction set includes a first direction, and the first direction is used to record the sound source direction of the sound that does not wake up the electronic device;
[0208] Step 904: If the sound source direction is in the first direction set, determine whether to wake up the electronic device according to the confidence level of the first direction corresponding to the sound source direction. The confidence level of the first direction is used to describe the probability that the voice for waking up the electronic device is emitted at the first direction.
[0209] In a possible implementation manner, it may further include:
[0210] If the sound source direction is not in the first direction set, extract the user voiceprint from the sound;
[0211] Determine whether the user voiceprint is included in the first voiceprint set; the first voiceprint set includes a first voiceprint, and the first voiceprint is used to record the user voiceprint of the sound that wakes up the electronic device;
[0212] Determine whether to wake up the electronic device according to the confidence level of the first voiceprint corresponding to the user voiceprint.
[0213] In a possible implementation manner, before calculating the sound source direction of the sound, it may further include:
[0214] Determine that the wake-up word confidence level is less than a second threshold; the second threshold is greater than the first threshold.
[0215] In a possible implementation manner, determining whether to wake up the electronic device according to the confidence level of the first direction corresponding to the sound source direction includes:
[0216] Determine whether the confidence level of the first azimuth corresponding to the sound source azimuth is less than the threshold a;
[0217] If it is less than the threshold a, the judgment result is not to wake up the electronic device;
[0218] If it is not less than the threshold a, the judgment result is to wake up the electronic device.
[0219] In a possible implementation manner, determining whether to wake up the electronic device according to the confidence level of the first voiceprint corresponding to the user's voiceprint may include:
[0220] Determine whether the confidence level of the first voiceprint corresponding to the user's voiceprint is less than the threshold c;
[0221] If it is less than the threshold c, the judgment result is not to wake up the electronic device;
[0222] If it is not less than the threshold c, the judgment result is to wake up the electronic device.
[0223] In a possible implementation manner, it may further include:
[0224] If the judgment result is to wake up the electronic device and the first voiceprint set includes the user's voiceprint of the sound, increase the confidence level of the first voiceprint corresponding to the user's voiceprint;
[0225] If the judgment result is to wake up the electronic device and the first voiceprint set does not include the user's voiceprint of the sound, store the user's voiceprint as the first voiceprint in the first voiceprint set and set an initial confidence level for the first voiceprint.
[0226] In a possible implementation manner, it may further include:
[0227] If the judgment result is not to wake up the electronic device and the first azimuth set includes the sound source azimuth of the sound, reduce the confidence level of the first azimuth including the sound source azimuth;
[0228] If the judgment result is not to wake up the electronic device and the first azimuth set does not include the sound source azimuth of the sound, store the sound source azimuth as the first azimuth in the first azimuth set and set an initial confidence level for the first azimuth.
[0229] Figure 9 The specific implementation of Figure 6 The embodiments shown may be referred to, and details are not described herein again.
[0230] Figure 10 is a flowchart of another embodiment of the wake-up method provided by the embodiments of the present application. This method can be applied to an electronic device such as the above-mentioned smart speaker. This method may include:
[0231] Step 1001: Receive a sound and calculate the confidence level of the wake-up word in the sound; the confidence level of the wake-up word is used to describe the probability that the sound includes the wake-up word sound.
[0232] Step 1002: If the confidence level of the wake-up word is greater than or equal to the first threshold, calculate the sound source direction of the sound.
[0233] Step 1003: Determine whether the sound source direction is in the second direction set; the second direction set includes a second direction, and the second direction is used to record the sound source direction of the sound that wakes up the electronic device.
[0234] Step 1004: If the sound source direction is in the second direction set, determine whether to wake up the electronic device according to the confidence level of the second direction corresponding to the sound source direction. The confidence level of the second direction is used to describe the probability that the voice that wakes up the electronic device is emitted at the second direction.
[0235] In a possible implementation, it may further include:
[0236] If the sound source direction is not in the second direction set, extract the user voiceprint from the sound.
[0237] Determine whether the user voiceprint is included in the first voiceprint set; the first voiceprint set includes a first voiceprint, and the first voiceprint is used to record the user voiceprint of the sound that wakes up the electronic device.
[0238] Determine whether to wake up the electronic device according to the confidence level of the first voiceprint corresponding to the user voiceprint.
[0239] In a possible implementation, before calculating the sound source direction of the sound, it may further include:
[0240] Determine that the confidence level of the wake-up word is less than the second threshold; the second threshold is greater than the first threshold.
[0241] In a possible implementation, determining whether to wake up the electronic device according to the confidence level of the second direction corresponding to the sound source direction may include:
[0242] Determine whether the confidence level of the second direction corresponding to the sound source direction is less than threshold b;
[0243] If it is less than threshold b, the determination result is not to wake up the electronic device;
[0244] If it is not less than threshold b, the determination result is to wake up the electronic device.
[0245] In a possible implementation, determining whether to wake up the electronic device according to the confidence level of the first voiceprint corresponding to the user voiceprint may include:
[0246] Determine whether the confidence level of the first voiceprint corresponding to the user voiceprint is less than threshold c;
[0247] If it is less than the threshold c, the judgment result is not to wake up the electronic device;
[0248] If it is not less than the threshold c, the judgment result is to wake up the electronic device.
[0249] In a possible implementation, it may further include:
[0250] If the judgment result is to wake up the electronic device and the sound source direction of the sound is included in the second direction set, increase the confidence of the second direction corresponding to the sound source direction;
[0251] If the judgment result is to wake up the electronic device and the sound source direction of the sound is not included in the second direction set, store the sound source direction as the second direction in the second direction set and set an initial confidence for the second direction.
[0252] In a possible implementation, it may further include:
[0253] If the judgment result is to wake up the electronic device and the user voiceprint of the sound is included in the first voiceprint set, increase the confidence of the first voiceprint corresponding to the user voiceprint;
[0254] If the judgment result is to wake up the electronic device and the user voiceprint of the sound is not included in the first voiceprint set, store the user voiceprint as the first voiceprint in the first voiceprint set and set an initial confidence for the first voiceprint.
[0255] Figure 10 The specific implementation of Figure 7 The embodiments shown may be referred to, and details are not described herein again.
[0256] It can be understood that some or all of the steps or operations in the above embodiments are only examples. The embodiments of the present application may also perform other operations or various deformations of the operations. In addition, each step may be executed in a different order presented in the above embodiments, and it is possible not to execute all the operations in the above embodiments.
[0257] Figure 11 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 11 shown, the electronic device 1100 may include: a calculation unit 1110 and a judgment unit 1120.
[0258] In one embodiment:
[0259] The calculation unit 1110 is configured to receive a sound and calculate the wake word confidence of the sound; the wake word confidence is used to describe the probability that the sound includes a wake word sound. If the wake word confidence is greater than or equal to the first threshold, calculate the sound source direction of the sound;
[0260] A determination unit 1120 is configured to determine whether the sound source direction is in the first direction set or the second direction set. The first direction set includes a first direction for recording the sound source direction of the sound that does not wake up the electronic device, and the second direction set includes a second direction for recording the sound source direction of the sound that wakes up the electronic device. If the sound source direction is only in the first direction set, it is determined whether to wake up the electronic device according to the confidence level of the first direction corresponding to the sound source direction, and the confidence level of the first direction is used to describe the probability of the voice for waking up the electronic device being emitted at the first direction. If the sound source direction is only in the second direction set, it is determined whether to wake up the electronic device according to the confidence level of the second direction corresponding to the sound source direction, and the confidence level of the second direction is used to describe the probability of the voice for waking up the electronic device being emitted at the second direction.
[0261] In a possible implementation, the determination unit 1120 may further be configured to: if the sound source direction is in both the first direction set and the second direction set, or the sound source direction is not in either the first direction set or the second direction set, extract the user voiceprint from the sound; determine whether the user voiceprint is included in the first voiceprint set; the first voiceprint set includes a first voiceprint for recording the user voiceprint of the sound that wakes up the electronic device; and determine whether to wake up the electronic device according to the confidence level of the first voiceprint corresponding to the user voiceprint.
[0262] In a possible implementation, the determination unit 1120 may further be configured to: before calculating the sound source direction of the sound, determine that the wake-up word confidence level is less than a second threshold, where the second threshold is greater than the first threshold.
[0263] In a possible implementation, the determination unit 1120 may specifically be configured to: determine whether the confidence level of the first direction corresponding to the sound source direction is less than a threshold a; if it is less than the threshold a, the determination result is not to wake up the electronic device; if it is not less than the threshold a, the determination result is to wake up the electronic device.
[0264] In a possible implementation, the determination unit 1120 may specifically be configured to: determine whether the confidence level of the second direction corresponding to the sound source direction is less than a threshold b; if it is less than the threshold b, the determination result is not to wake up the electronic device; if it is not less than the threshold b, the determination result is to wake up the electronic device.
[0265] In a possible implementation, the determination unit 1120 may specifically be configured to: determine whether the confidence level of the first voiceprint corresponding to the user voiceprint is less than a threshold c; if it is less than the threshold c, the determination result is not to wake up the electronic device; if it is not less than the threshold c, the determination result is to wake up the electronic device.
[0266] In a possible implementation, it may further include: an updating unit, configured to, if the judgment result is to wake up the electronic device and the second azimuth set includes the sound source azimuth of the sound, increase the confidence level of the second azimuth corresponding to the sound source azimuth; if the judgment result is to wake up the electronic device and the second azimuth set does not include the sound source azimuth of the sound, store the sound source azimuth as the second azimuth into the second azimuth set, and set an initial confidence level for the second azimuth.
[0267] In a possible implementation, the updating unit may further be configured to: if the judgment result is to wake up the electronic device and the first voiceprint set includes the user voiceprint of the sound, increase the confidence level of the first voiceprint corresponding to the user voiceprint; if the judgment result is to wake up the electronic device and the first voiceprint set does not include the user voiceprint of the sound, store the user voiceprint as the first voiceprint into the first voiceprint set, and set an initial confidence level for the first voiceprint.
[0268] In a possible implementation, the updating unit may further be configured to: if the judgment result is not to wake up the electronic device and the first azimuth set includes the sound source azimuth of the sound, decrease the confidence level of the first azimuth including the sound source azimuth; if the judgment result is not to wake up the electronic device and the first azimuth set does not include the sound source azimuth of the sound, store the sound source azimuth as the first azimuth into the first azimuth set, and set an initial confidence level for the first azimuth.
[0269] In another embodiment:
[0270] A calculation unit 1110, configured to receive a sound; calculate the wake-up word confidence level of the sound; the wake-up word confidence level is used to describe the probability that the sound includes the wake-up word sound; if the wake-up word confidence level is greater than or equal to a first threshold, calculate the sound source azimuth of the sound;
[0271] A judgment unit 1120, configured to judge whether the sound source azimuth is in the first azimuth set; wherein, the first azimuth set includes a first azimuth, and the first azimuth is used to record the sound source azimuth of the sound that does not wake up the electronic device; if the sound source azimuth is in the first azimuth set, judge whether to wake up the electronic device according to the confidence level of the first azimuth corresponding to the sound source azimuth, and the confidence level of the first azimuth is used to describe the probability that the voice for waking up the electronic device is emitted at the first azimuth.
[0272] In a possible implementation, the judgment unit 1120 may further be configured to: if the sound source azimuth is not in the first azimuth set, extract the user voiceprint according to the sound; judge whether the first voiceprint set includes the user voiceprint; the first voiceprint set includes a first voiceprint, and the first voiceprint is used to record the user voiceprint of the sound that wakes up the electronic device; judge whether to wake up the electronic device according to the confidence level of the first voiceprint corresponding to the user voiceprint.
[0273] In a possible implementation, the determination unit 1120 can also be configured to: before calculating the sound source direction of the sound, determine that the wake word confidence is less than a second threshold; the second threshold is greater than the first threshold.
[0274] In a possible implementation, the determination unit 1120 can specifically be configured to: determine whether the confidence of the first direction corresponding to the sound source direction is less than a threshold a; if it is less than the threshold a, the determination result is not to wake up the electronic device; if it is not less than the threshold a, the determination result is to wake up the electronic device.
[0275] In a possible implementation, the determination unit 1120 can specifically be configured to: determine whether the confidence of the first voiceprint corresponding to the user voiceprint is less than a threshold c; if it is less than the threshold c, the determination result is not to wake up the electronic device; if it is not less than the threshold c, the determination result is to wake up the electronic device.
[0276] In a possible implementation, it may further include:
[0277] An update unit, configured to, if the determination result is to wake up the electronic device and the first voiceprint set includes the user voiceprint of the sound, increase the confidence of the first voiceprint corresponding to the user voiceprint; if the determination result is to wake up the electronic device and the first voiceprint set does not include the user voiceprint of the sound, store the user voiceprint as the first voiceprint in the first voiceprint set and set an initial confidence for the first voiceprint.
[0278] In a possible implementation, the update unit can also be configured to: if the determination result is not to wake up the electronic device and the first direction set includes the sound source direction of the sound, decrease the confidence of the first direction including the sound source direction; if the determination result is not to wake up the electronic device and the first direction set does not include the sound source direction of the sound, store the sound source direction as the first direction in the first direction set and set an initial confidence for the first direction.
[0279] In yet another embodiment:
[0280] A calculation unit 1110, configured to receive a sound; calculate the wake word confidence of the sound; the wake word confidence is used to describe the probability that the sound includes a wake word sound; if the wake word confidence is greater than or equal to the first threshold, calculate the sound source direction of the sound;
[0281] A determination unit 1120, configured to determine whether the sound source direction is in a second direction set; wherein, the second direction set includes a second direction, and the second direction is used to record the sound source direction of the sound for waking up the electronic device; if the sound source direction is in the second direction set, determine whether to wake up the electronic device according to the confidence of the second direction corresponding to the sound source direction, and the confidence of the second direction is used to describe the probability that a voice for waking up the electronic device is emitted at the second direction.
[0282] In a possible implementation, the determination unit 1120 may further be configured to: if the sound source direction is not in the second direction set, extract the user's voiceprint from the sound; determine whether the first voiceprint set includes the user's voiceprint; the first voiceprint set includes a first voiceprint, and the first voiceprint is used to record the user's voiceprint of the sound for waking up the electronic device; determine whether to wake up the electronic device according to the confidence level of the first voiceprint corresponding to the user's voiceprint.
[0283] In a possible implementation, the determination unit 1120 may further be configured to: before calculating the sound source direction of the sound, determine that the wake-up word confidence level is less than a second threshold; the second threshold is greater than the first threshold.
[0284] In a possible implementation, the determination unit 1120 may specifically be configured to: determine whether the confidence level of the second direction corresponding to the sound source direction is less than a threshold b; if it is less than the threshold b, the determination result is not to wake up the electronic device; if it is not less than the threshold b, the determination result is to wake up the electronic device.
[0285] In a possible implementation, the determination unit 1120 may specifically be configured to: determine whether the confidence level of the first voiceprint corresponding to the user's voiceprint is less than a threshold c; if it is less than the threshold c, the determination result is not to wake up the electronic device; if it is not less than the threshold c, the determination result is to wake up the electronic device.
[0286] In a possible implementation, it may further include:
[0287] An update unit, configured to, if the determination result is to wake up the electronic device and the second direction set includes the sound source direction of the sound, increase the confidence level of the second direction corresponding to the sound source direction; if the determination result is to wake up the electronic device and the second direction set does not include the sound source direction of the sound, store the sound source direction as the second direction in the second direction set and set an initial confidence level for the second direction.
[0288] In a possible implementation, the update unit may further be configured to: if the determination result is to wake up the electronic device and the first voiceprint set includes the user's voiceprint of the sound, increase the confidence level of the first voiceprint corresponding to the user's voiceprint; if the determination result is to wake up the electronic device and the first voiceprint set does not include the user's voiceprint of the sound, store the user's voiceprint as the first voiceprint in the first voiceprint set and set an initial confidence level for the first voiceprint.
[0289] Figure 11 The electronic device provided in the illustrated embodiment may be used to execute the technical solution of the method embodiment of the present application Figures 5 - 7 The implementation principle and technical effects can be further referred to the relevant descriptions in the method embodiment.
[0290] It should be understood that the above Figure 11The division of each unit of the device shown is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these units can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; or some units can be implemented in the form of software called by a processing element, and some units can be implemented in the form of hardware. For example, the acquisition unit can be a separately established processing element, or can be integrated in a certain chip of the electronic device. The implementation of other units is similar. In addition, these units can be fully or partially integrated together, or can be independently implemented. In the implementation process, each step of the above method or each of the above units can be completed by the integrated logic circuit in the processor element or the instructions in the form of software.
[0291] The embodiments of the present application also provide an electronic device, including: a processor; a memory; and a computer program, where the computer program is stored in the memory, and the computer program includes instructions, when the instructions are executed by the device, the device is caused to execute Figures 5 - 7 the method shown.
[0292] The embodiments of the present application also provide a computer-readable storage medium, in which a computer program is stored, and when it runs on a computer, the computer is caused to execute the present application Figures 5 - 7 the method provided by the shown embodiment.
[0293] The embodiments of the present application also provide a computer program product, which includes a computer program, and when it runs on a computer, the computer is caused to execute the present application Figures 5 - 7 the method provided by the shown embodiment.
[0294] Those of ordinary skill in the art can realize that the units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0295] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0296] In several embodiments provided by the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0297] As described above, the foregoing are only specific implementation manners of the present application. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all such changes or substitutions should be covered by the protection scope of the present application. The protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A wake-up method, applied to an electronic device including a pickup and a speaker, the pickup including a plurality of microphones, characterized in that, the method includes: receiving sound; calculating a wake-word confidence level of the sound; the wake-word confidence level is used to represent the probability that the sound includes a wake word; after the wake-word confidence level is greater than or equal to a first threshold, calculating a sound source azimuth of the sound; after the sound source azimuth matches a first azimuth in a first azimuth set, and after a first azimuth confidence level corresponding to the matched first azimuth is greater than or equal to a third threshold, waking up the electronic device; or, after the first azimuth confidence level corresponding to the matched first azimuth is less than the third threshold, not waking up the electronic device; wherein, the wake word is used to wake up the electronic device; the sound source azimuth is the direction and position of the sound source relative to the electronic device; the first azimuth set includes M first azimuth elements, each first azimuth element including a first azimuth and a first azimuth confidence level; the first azimuth is the direction and position of the sound source for waking up the electronic device relative to the electronic device; the first azimuth confidence level is used to represent the probability of waking up the electronic device at the first azimuth; M is a positive integer greater than or equal to 1; the method further includes: when the electronic device is woken up and the sound source azimuth of the sound for waking up the electronic device is not included in the first azimuth set, adding the sound source azimuth of the sound for waking up the electronic device as a first azimuth to the first azimuth set.
2. The method according to claim 1, characterized in that, the calculating the azimuth of the sound source corresponding to the sound after the wake-word confidence level is greater than or equal to the first threshold; includes: after the wake-word confidence level is greater than or equal to the first threshold and less than a second threshold, calculating the azimuth of the sound source corresponding to the sound.
3. The method according to claim 1 or 2, characterized in that, the matching of the sound source azimuth with a first azimuth in the first azimuth set; includes: the angular deviation between the direction of the sound source azimuth relative to the electronic device and the direction of a first azimuth in the first azimuth set relative to the electronic device is within a preset fourth threshold; and, the position deviation between the position of the sound source azimuth relative to the electronic device and the position of the first azimuth relative to the electronic device is within a preset fifth threshold.
4. The method according to any one of claims 1 or 2, characterized in that, the method further includes: after the sound source azimuth does not match any first azimuth in the first azimuth set, then extracting a voiceprint from the sound; after the voiceprint matches a first voiceprint in a first voiceprint set, and after a first voiceprint confidence level corresponding to the first voiceprint is greater than or equal to a preset sixth threshold, waking up the electronic device; or, after the first voiceprint confidence level corresponding to the first voiceprint is less than the preset sixth threshold, not waking up the electronic device; Wherein, the first voiceprint set includes L voiceprint elements, each voiceprint element includes a first voiceprint and a first voiceprint confidence level. The first voiceprint is used to represent the voiceprint for waking up the electronic device, and the first voiceprint confidence level is used to represent the probability that the first voiceprint wakes up the electronic device; L is a positive integer greater than or equal to 1.
5. The method according to claim 4, wherein, after waking up the electronic device, the method further includes: updating the first voiceprint set.
6. The method according to any one of claim 4, wherein, after the wake-up word confidence level is greater than or equal to a second threshold, wake up the electronic device, and update the first azimuth set and the first voiceprint set.
7. A wake-up method applied to an electronic device including a pickup and a speaker, the pickup includes a plurality of microphones, wherein, the method includes: receiving a sound; calculating the wake-up word confidence level of the sound; the wake-up word confidence level is used to represent the probability that the sound includes a wake-up word; after the wake-up word confidence level is greater than or equal to a first threshold, calculating the sound source azimuth of the sound; after the sound source azimuth matches a second azimuth in the second azimuth set, and after the second azimuth confidence level corresponding to the matched second azimuth is greater than or equal to a seventh threshold, wake up the electronic device; or, after the second azimuth confidence level corresponding to the matched second azimuth is less than the seventh threshold, do not wake up the electronic device; wherein, the wake-up word is used to wake up the electronic device; the sound source azimuth is the direction and position of the sound source relative to the electronic device; the second azimuth set includes N second azimuth elements, each second azimuth element includes a second azimuth and a second azimuth confidence level. The second azimuth is the direction and position of the sound source that has not woken up the electronic device relative to the electronic device, and the second azimuth confidence level is used to represent the probability that the electronic device is not woken up at the second azimuth; N is a positive integer greater than or equal to 1; the method further includes: when the electronic device is not woken up and the second azimuth set does not include the sound source azimuth of the sound that has not woken up the electronic device, adding the sound source azimuth of the sound that has not woken up the electronic device as a second azimuth to the second azimuth set.
8. The method according to claim 7, wherein, the sound source azimuth matches a second azimuth in the second azimuth set; including: the angular deviation between the direction of the sound source azimuth relative to the electronic device and the direction of a second azimuth in the second azimuth set relative to the electronic device is within a preset eighth threshold; and, the position deviation between the position of the sound source azimuth relative to the electronic device and the position of the second azimuth relative to the electronic device is within a preset ninth threshold.
9. The method according to claim 7 or 8, wherein, the method further includes: after the sound source azimuth does not match any second azimuth in the second azimuth set, then extracting the voiceprint from the sound; After the voiceprint does not match any of the first voiceprints in the first voiceprint set, update the second azimuth set; Among them, the first voiceprint set includes L voiceprint elements, each voiceprint element includes a first voiceprint and a first voiceprint confidence level. The first voiceprint is used to represent the voiceprint for waking up the electronic device, and the first voiceprint confidence level is used to represent the probability that the first voiceprint wakes up the electronic device; L is a positive integer greater than or equal to 1.
10. The method according to claim 7 or 8, characterized in that, The method further includes: After the sound source azimuth does not match any of the second azimuths in the second azimuth set, then Extract the voiceprint from the sound; After the voiceprint matches one of the first voiceprints in the first voiceprint set, and After the first voiceprint confidence level corresponding to the first voiceprint is greater than or equal to a preset tenth threshold, wake up the electronic device; or, After the first voiceprint confidence level corresponding to the first voiceprint is less than the preset tenth threshold, do not wake up the electronic device and update the second azimuth set; Among them, the first voiceprint set includes L voiceprint elements, each voiceprint element includes a first voiceprint and a first voiceprint confidence level. The first voiceprint confidence level is used to represent the probability that the first voiceprint wakes up the electronic device, and the first voiceprint is used to represent the voiceprint for waking up the electronic device; L is a positive integer greater than or equal to 1.
11. The method according to claim 10, characterized in that, After waking up the electronic device, the method further includes: updating the first voiceprint set.
12. The method according to claim 10, characterized in that, After the wake-up word confidence level is greater than or equal to the second threshold, wake up the electronic device and update the first voiceprint set.
13. A wake-up method applied to an electronic device including a pickup and a speaker, the pickup includes a plurality of microphones, characterized in that, The method includes: Receiving sound; Calculating the wake-up word confidence level of the sound; the wake-up word confidence level is used to represent the probability that the sound includes a wake-up word; After the wake-up word confidence level is greater than or equal to the first threshold, calculate the sound source azimuth of the sound; After the sound source azimuth matches one of the second azimuths in the second azimuth set, and after the sound source azimuth does not match any of the first azimuths in the first azimuth set, and After the second azimuth confidence level corresponding to the matched second azimuth is greater than or equal to the eleventh threshold, wake up the electronic device; or, After the second azimuth confidence level corresponding to the matched second azimuth is less than the eleventh threshold, do not wake up the electronic device; Wherein, the wake word is used to wake up the electronic device; the sound source orientation is the direction and position of the sound source relative to the electronic device; the first orientation set includes M first orientation elements, and each first orientation element includes a first orientation and a first orientation confidence; the first orientation is the direction and position of the sound source that wakes up the electronic device relative to the electronic device; the first orientation confidence is used to represent the probability of waking up the electronic device in the first orientation; the second orientation set includes N second orientation elements, and each second orientation element includes a second orientation and a second orientation confidence; the second orientation is the direction and position of the sound source that does not wake up the electronic device relative to the electronic device; the second orientation confidence is used to represent the confidence of not waking up the electronic device in the second orientation; both M and N are positive integers greater than or equal to 1. The method further includes: When the electronic device is woken up and the sound source orientation of the sound that wakes up the electronic device is not included in the first orientation set, adding the sound source orientation of the sound that wakes up the electronic device as a first orientation to the first orientation set; and, the electronic device is used to add the sound source orientation of the sound that does not wake up the electronic device as a second orientation to the second orientation set when the electronic device is not woken up and the sound source orientation of the sound that does not wake up the electronic device is not included in the second orientation set.
14. The method according to claim 13, characterized in that the sound source orientation matches one of the second orientations in the second orientation set; including: the angle deviation between the direction of the sound source orientation relative to the electronic device and the direction of one of the second orientations in the second orientation set relative to the electronic device is within a preset twelfth threshold; and, the position deviation between the position of the sound source orientation relative to the electronic device and the position of the second orientation relative to the electronic device is within a preset thirteenth threshold; the sound source orientation does not match any of the first orientations in the first orientation set; including: the angle deviation between the direction of the sound source orientation relative to the electronic device and the direction of any of the first orientations in the first orientation set relative to the electronic device is not within a preset fourteenth threshold; and, the position deviation between the position of the sound source orientation relative to the electronic device and the position of any of the first orientations in the first orientation set relative to the electronic device is not within a preset fifteenth threshold.
15. The method according to claim 13, characterized in that The method further includes: after the sound source orientation matches one of the first orientations in the first orientation set, and after the sound source orientation does not match any of the second orientations in the second orientation set, and after the first orientation confidence corresponding to the matched first orientation is greater than or equal to a sixteenth threshold, then wake up the electronic device; or, after the first orientation confidence corresponding to the matched first orientation is less than the sixteenth threshold, then do not wake up the electronic device.
16. The method according to claim 15, characterized in that The sound source orientation matches one of the first orientations in the first orientation set, including: The angle deviation between the direction of the sound source orientation relative to the electronic device and the direction of one of the first orientations in the first orientation set relative to the electronic device is within a preset fourteenth threshold; and The position deviation between the position of the sound source orientation relative to the electronic device and the position of the first orientation relative to the electronic device is within a preset fifteenth threshold; The sound source orientation does not match any of the second orientations in the second orientation set, including: The angle deviation between the direction of the sound source orientation relative to the electronic device and the direction of any of the second orientations in the second orientation set relative to the electronic device is not within a preset twelfth threshold; and The position deviation between the position of the sound source orientation relative to the electronic device and the position of any of the second orientations in the second orientation set relative to the electronic device is not within a preset thirteenth threshold.
17. The method according to claim 13, wherein, The method further includes: After the sound source orientation does not match any of the second orientations in the second orientation set, and after the sound source orientation does not match any of the first orientations in the first orientation set, then Extract a voiceprint from the sound; After the voiceprint matches one of the first voiceprints in the first voiceprint set, and After the first voiceprint confidence corresponding to the first voiceprint is greater than or equal to a preset sixteenth threshold, then wake up the electronic device and update the first orientation combination and the first voiceprint set; or, After the first voiceprint confidence corresponding to the first voiceprint is less than the preset sixteenth threshold, then do not wake up the electronic device and update the second orientation set; wherein, the first voiceprint set includes L voiceprint elements, each voiceprint element includes a first voiceprint and a first voiceprint confidence, the first voiceprint confidence is used to represent the probability that the first voiceprint wakes up the electronic device, and the first voiceprint is used to represent the voiceprint for waking up the electronic device; L is a positive integer greater than or equal to 1.
18. The method according to claim 17, wherein, The method further includes: After the voiceprint does not match any of the first voiceprints in the first voiceprint set, update the second orientation set.
19. An electronic device, including a pickup and a speaker, the pickup includes a plurality of microphones, wherein, The electronic device further includes: A processor; A memory; And a computer program, wherein the computer program is stored in the memory, and when the computer program is executed by the processor, the electronic device executes the method according to any one of claims 1-18.
20. A computer-readable storage medium, wherein, The computer-readable storage medium includes a computer program, and when the computer program runs on an electronic device, the electronic device executes the method according to any one of claims 1-18, wherein the electronic device includes a pickup and a speaker, and the pickup includes a plurality of microphones.
Citation Information
Patent Citations
Equipment control method and device, storage medium and electronic device
CN108800473A
Voice wake-up method and system for intelligent equipment
CN109461449A