Voice wake-up method, apparatus, device, storage medium, and program product
By acquiring bone conduction signals with a bone conduction microphone and fusing them with air conduction signals, and using an acoustic model to process the posterior probability vector for wake-up word detection, the problem of low wake-up word recognition accuracy in voice wake-up devices is solved, achieving high-precision voice wake-up.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-30
- Publication Date
- 2026-04-10
AI Technical Summary
In existing technologies, voice wake-up devices suffer from low accuracy in wake-up word recognition and low voice wake-up accuracy due to the loss of air conduction signals caused by the delay in voice detection of bone conduction microphones.
Speech detection is performed by collecting bone conduction signals through a bone conduction microphone. The bone conduction signals and air conduction signals are combined for signal fusion or head loss compensation. The posterior probability vector is processed using an acoustic model to detect wake words, thereby improving the accuracy of wake word recognition.
It improves the accuracy of wake word recognition and voice wake-up, while ensuring low power consumption and improving the accuracy of wake word detection.
Smart Images

Figure CN115731927B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of speech recognition, and particularly relate to a method and apparatus for voice wake-up, a device, a storage medium and a program product. BACKGROUND
[0002] Currently, more and more intelligent devices complete tasks through voice control. Generally, an intelligent device needs to be woken up by a user voice input of a wake-up word, so as to receive an instruction to complete a task. In addition, with the development of bone conduction devices, a large number of bone conduction microphones are applied to wearable devices to wake up intelligent devices through wearable devices. The wearable devices are, for example, wireless earphones, smart glasses, smart watches, etc. Among them, the sensor in the bone conduction microphone is a non-acoustic sensor, which converts the vibration signal of the vocal cords when people speak into an electric signal, and the electric signal is called a bone conduction signal.
[0003] In the related art, a wearable device is provided with a bone conduction microphone and an air microphone. In order to realize low power consumption of the wearable device, the air microphone is in a dormant state before the intelligent device is woken up. Since the bone conduction microphone has low power consumption, the bone conduction microphone can be used to collect a bone conduction signal, voice detection (such as voice activate detector (VAD)) is performed based on the bone conduction signal, and the air microphone is controlled to be turned on or off to reduce power consumption. In the case where it is confirmed through voice detection that there is a voice input, the air microphone is turned on, an air conduction signal is collected through the air microphone, and wake-up word recognition is performed based on the air conduction signal, that is, voice wake-up is performed.
[0004] However, since there is an algorithm delay in voice detection, the voice head of the input command word may be truncated, that is, the collected air conduction signal may be truncated, so that the complete information of the command word input by the sound source is not included, resulting in low recognition accuracy of the wake-up word and low accuracy of voice wake-up. SUMMARY
[0005] Embodiments of the present application provide a method and apparatus for voice wake-up, a device, a storage medium and a program product, which can improve the accuracy of voice wake-up. The technical solution is as follows:
[0006] In a first aspect, a method for voice wake-up is provided, which comprises:
[0007] Voice detection is performed according to a bone conduction signal collected by a bone conduction microphone, the bone conduction signal containing command word information input by a sound source; in the case where it is detected that there is a voice input, wake-up word detection is performed based on the bone conduction signal; and in the case where it is detected that the command word includes a wake-up word, a device to be woken up is voice-woken up.
[0008] In the embodiment of the present application, the bone conduction microphone is used to collect the bone conduction signal for voice detection, which can ensure low power consumption. In addition, considering that the delay of voice detection may cause the collected air conduction signal to lose the head, thereby not containing the complete information of the command word input by the sound source, the bone conduction signal collected by the bone conduction microphone contains the command word information input by the sound source, that is, the bone conduction signal does not lose the head, so the present scheme detects the wake-up word based on the bone conduction signal. In this way, the recognition accuracy of the wake-up word is high, and the accuracy of voice wake-up is high.
[0009] Optionally, the detection of the wake-up word based on the bone conduction signal comprises: determining a fusion signal based on the bone conduction signal; and detecting the wake-up word based on the fusion signal. It should be noted that the fusion signal determined based on the bone conduction signal also contains the command word information input by the sound source.
[0010] Optionally, before the determination of the fusion signal based on the bone conduction signal, the method further comprises: turning on the air microphone, and collecting an air conduction signal through the air microphone; and the determination of the fusion signal based on the bone conduction signal comprises: fusing a starting part of the bone conduction signal and the air conduction signal to obtain the fusion signal, the starting part of the bone conduction signal being determined according to a detection delay of the voice detection; or generating an enhanced starting signal based on the starting part of the bone conduction signal, and fusing the enhanced starting signal and the air conduction signal to obtain the fusion signal, the starting part of the bone conduction signal being determined according to the detection delay of the voice detection; or directly fusing the bone conduction signal and the air conduction signal to obtain the fusion signal. That is, the present embodiment provides three methods for compensating for the loss of the head of the air conduction signal by the bone conduction signal, that is, directly fusing the signals to compensate for the loss of the head of the air conduction signal. Optionally, the signal fusion is performed by signal splicing.
[0011] Optionally, the determination of the fusion signal based on the bone conduction signal comprises: determining the bone conduction signal as the fusion signal. That is, the present embodiment can also directly use the bone conduction signal to detect the wake-up word.
[0012] Optionally, the detection of the wake-up word based on the fusion signal comprises: inputting a plurality of audio frames included in the fusion signal into a first acoustic model to obtain a plurality of posterior probability vectors output by the first acoustic model, the plurality of posterior probability vectors corresponding to the plurality of audio frames one by one, a first posterior probability vector in the plurality of posterior probability vectors being used for a probability that a phoneme of a first audio frame in the plurality of audio frames belongs to a plurality of specified phonemes; and detecting the wake-up word based on the plurality of posterior probability vectors. That is, the fusion signal is first processed by the first acoustic model to obtain a plurality of posterior probability vectors corresponding to a plurality of audio frames included in the fusion signal, and then the wake-up word is detected based on the plurality of posterior probability vectors, for example, the plurality of posterior probability vectors are decoded to detect the wake-up word.
[0013] Optionally, before the detecting the wake-up word based on the bone conduction signal, the method further comprises: turning on an air microphone, and collecting an air conduction signal through the air microphone; and the detecting the wake-up word based on the bone conduction signal comprises: determining a plurality of posterior probability vectors based on the bone conduction signal and the air conduction signal, the plurality of posterior probability vectors corresponding to a plurality of audio frames included in the bone conduction signal and the air conduction signal one by one, a first posterior probability vector in the plurality of posterior probability vectors being used to indicate a probability that a phoneme in a first audio frame in the plurality of audio frames belongs to a plurality of specified phonemes; and detecting the wake-up word based on the plurality of posterior probability vectors. That is, in the embodiments of the present application, signal fusion can also not be performed, but the posterior probability vector corresponding to each audio frame can be directly determined based on the bone conduction signal and the air conduction signal, so that the plurality of posterior probability vectors obtained implicitly contain the command word information of the sound source input in the form of phoneme probability, that is, the bone conduction signal implicitly compensates for the loss of the air conduction signal.
[0014] Optionally, the determining the plurality of posterior probability vectors based on the bone conduction signal and the air conduction signal comprises: inputting a starting portion of the bone conduction signal and the air conduction signal into a second acoustic model to obtain a first number of bone conduction posterior probability vectors and a second number of air conduction posterior probability vectors output by the second acoustic model, the starting portion of the bone conduction signal being determined according to a detection time delay of voice detection, the first number of bone conduction posterior probability vectors corresponding to audio frames included in the starting portion of the bone conduction signal one by one, and the second number of air conduction posterior probability vectors corresponding to audio frames included in the air conduction signal one by one; fusing a first bone conduction posterior probability vector and a first air conduction posterior probability vector to obtain a second posterior probability vector, the first bone conduction posterior probability vector corresponding to a last audio frame of the starting portion of the bone conduction signal, a time length of the last audio frame being less than a frame time length, the first air conduction posterior probability vector corresponding to a first audio frame of the air conduction signal, a time length of the first audio frame being less than the frame time length, and the plurality of posterior probability vectors including the second posterior probability vector, vectors other than the first bone conduction posterior probability vector in the first number of bone conduction posterior probability vectors, and vectors other than the first air conduction posterior probability vector in the second number of air conduction posterior probability vectors. That is, in the embodiments of the present application, the starting portion of the bone conduction signal and the air conduction signal can be processed by the second acoustic model respectively to obtain the corresponding bone conduction posterior probability vectors and air conduction posterior probability vectors, and then the first bone conduction posterior probability vector and the first air conduction posterior probability vector are fused to implicitly compensate for the loss of the air conduction signal by the bone conduction signal.
[0015] Optionally, the multiple posterior probability vectors are determined based on the bone conduction signal and the air conduction signal, including: inputting a starting portion of the bone conduction signal and the air conduction signal into a third acoustic model to obtain the multiple posterior probability vectors output by the third acoustic model, the starting portion of the bone conduction signal being determined according to a detection time delay of the speech detection; or inputting the bone conduction signal and the air conduction signal into the third acoustic model to obtain the multiple posterior probability vectors output by the third acoustic model. That is, in the embodiments of the present application, the starting portion of the bone conduction signal and the air conduction signal can be input into the third acoustic model respectively, and the multiple posterior probability vectors are directly obtained through the third acoustic model. That is, the two signals are implicitly fused in the process of processing the starting portion of the bone conduction signal and the air conduction signal in the third acoustic model, that is, the bone conduction signal is implicitly used to perform head loss compensation on the air conduction signal.
[0016] Optionally, the detection of the wake-up word based on the multiple posterior probability vectors includes: determining a confidence that the phoneme sequence corresponding to the command word includes the phoneme sequence corresponding to the wake-up word based on the multiple posterior probability vectors and the phoneme sequence corresponding to the wake-up word; and determining that the command word includes the wake-up word in a case where the confidence exceeds a confidence threshold. For example, the confidence is obtained by decoding the multiple posterior probability vectors, and then the confidence threshold is used to determine whether the command word includes the wake-up word, that is, in a case where the confidence condition is met, it is determined that the command word contains the wake-up word.
[0017] Optionally, the detection of the wake-up word based on the multiple posterior probability vectors includes: determining a confidence that the phoneme sequence corresponding to the command word includes the phoneme sequence corresponding to the wake-up word based on the multiple posterior probability vectors and the phoneme sequence corresponding to the wake-up word; and determining that the command word includes the wake-up word in a case where the confidence exceeds a confidence threshold and a distance condition between the multiple posterior probability vectors and multiple template vectors is met, the multiple template vectors indicating probabilities that phonemes of a speech signal containing complete information of the wake-up word belong to multiple specified phonemes. That is, in a case where the confidence condition and the template matching are met, it is determined that the command word contains the wake-up word, so as to avoid false wake-up as much as possible.
[0018] Optionally, in a case where the multiple posterior probability vectors and the multiple template vectors correspond to each other one by one, the distance condition includes: a mean value of distances between the multiple posterior probability vectors and corresponding template vectors is less than a distance threshold. That is, whether the template is matched can be determined by the average distance between vectors.
[0019] Optionally, the method further comprises: obtaining a bone conduction registration signal, the bone conduction registration signal containing complete information of the wake-up word; and determining the confidence threshold and the plurality of template vectors based on the bone conduction registration signal and a phoneme sequence corresponding to the wake-up word. That is, in the registration process of the wake-up word, the confidence threshold and the plurality of template vectors can also be determined based on the bone conduction registration signal containing complete information of the wake-up word. The wake-up word detection in the subsequent voice wake-up process is performed by using the confidence threshold and the plurality of template vectors obtained in this way, which can improve the accuracy of the wake-up word detection and thus reduce false wake-up.
[0020] Optionally, the determining of the confidence threshold and the plurality of template vectors based on the bone conduction registration signal and the phoneme sequence corresponding to the wake-up word comprises: determining a fusion registration signal based on the bone conduction registration signal; and determining the confidence threshold and the plurality of template vectors based on the fusion registration signal and the phoneme sequence corresponding to the wake-up word. That is, in the registration process of the wake-up word, the fusion registration signal can also be obtained by signal fusion first. The fusion registration signal contains information of the command word input by the sound source, and the confidence threshold and the plurality of template vectors are then determined based on the fusion registration signal.
[0021] Optionally, the determining of the confidence threshold and the plurality of template vectors based on the fusion registration signal and the phoneme sequence corresponding to the wake-up word comprises: inputting a plurality of registration audio frames included in the fusion registration signal into a first acoustic model to obtain a plurality of registration posterior probability vectors output by the first acoustic model, the plurality of registration posterior probability vectors corresponding to the plurality of registration audio frames one by one, a first registration posterior probability vector in the plurality of registration posterior probability vectors indicating a probability that a phoneme of a first registration audio frame in the plurality of registration audio frames belongs to a plurality of specified phonemes; determining the plurality of registration posterior probability vectors as the plurality of template vectors; and determining the confidence threshold based on the plurality of registration posterior probability vectors and the phoneme sequence corresponding to the wake-up word. That is, similar to the processing of the fusion signal in the voice wake-up process, in the registration process of the wake-up word, the fusion registration signal can also be processed by the first acoustic model first to obtain a plurality of registration posterior probability vectors corresponding to a plurality of registration audio frames included in the fusion registration signal, and then the confidence threshold is determined based on the plurality of registration posterior probability vectors and the phoneme sequence corresponding to the wake-up word. For example, the plurality of registration posterior probability vectors are decoded to determine the confidence threshold. In addition, the plurality of registration posterior probability vectors can also be determined as the plurality of template vectors.
[0022] Optionally, before determining the confidence threshold based on the bone conduction registration signal and the phoneme sequence corresponding to the wake-up word, the method further includes: obtaining an air conduction registration signal; and determining the confidence threshold based on the bone conduction registration signal and the phoneme sequence corresponding to the wake-up word includes: determining a plurality of registration posterior probability vectors based on the bone conduction registration signal and the air conduction registration signal, the plurality of registration posterior probability vectors corresponding one-to-one to a plurality of registration audio frames included in the bone conduction registration signal and the air conduction registration signal, a first registration posterior probability vector in the plurality of registration posterior probability vectors indicating a probability that a phoneme of a first registration audio frame in the plurality of registration audio frames belongs to the plurality of specified phonemes; and determining the confidence threshold based on the plurality of registration posterior probability vectors and the phoneme sequence corresponding to the wake-up word. That is, in the registration process of the wake-up word, signal fusion can not be performed first, but the registration posterior probability vector corresponding to each registration audio frame can be determined based on the bone conduction registration signal and the air conduction registration signal.
[0023] In a second aspect, a voice wake-up apparatus is provided, which has functions to implement the behaviors of the voice wake-up method in the first aspect. The voice wake-up apparatus includes one or more modules for implementing the voice wake-up method provided in the first aspect.
[0024] That is, a voice wake-up apparatus is provided, which includes:
[0025] a voice detection module configured to perform voice detection based on a bone conduction signal collected by a bone conduction microphone, the bone conduction signal containing command word information input by a sound source;
[0026] a wake-up word detection module configured to, when detecting that there is voice input, perform detection of a wake-up word based on the bone conduction signal;
[0027] a voice wake-up module configured to, when detecting that the command word includes the wake-up word, perform voice wake-up on a device to be woken up.
[0028] Optionally, the wake-up word detection module includes:
[0029] a first determination sub-module configured to determine a fusion signal based on the bone conduction signal;
[0030] a wake-up word detection sub-module configured to perform detection of the wake-up word on the fusion signal.
[0031] Optionally, the apparatus further includes:
[0032] a processing module configured to turn on an air microphone and collect an air conduction signal by the air microphone;
[0033] the first determination sub-module is configured to:
[0034] fusing a starting portion of a bone conduction signal and an air conduction signal to obtain a fused signal, the starting portion of the bone conduction signal being determined according to a detection time delay of the speech detection; or
[0035] generating an enhanced starting signal based on a starting portion of a bone conduction signal, fusing the enhanced starting signal and an air conduction signal to obtain a fused signal, the starting portion of the bone conduction signal being determined according to a detection time delay of the speech detection; or
[0036] directly fusing a bone conduction signal and an air conduction signal to obtain a fused signal.
[0037] Optionally, the wake-up word detection sub-module is configured to:
[0038] inputting a plurality of audio frames included in the fused signal into the first acoustic model to obtain a plurality of posterior probability vectors output by the first acoustic model, the plurality of posterior probability vectors corresponding to the plurality of audio frames in a one-to-one manner, a first posterior probability vector in the plurality of posterior probability vectors being used to indicate a probability that a phoneme of a first audio frame in the plurality of audio frames belongs to a plurality of specified phonemes;
[0039] detecting the wake-up word based on the plurality of posterior probability vectors.
[0040] Optionally, the apparatus further comprises:
[0041] the processing module is configured to turn on the air microphone and collect the air conduction signal through the air microphone;
[0042] the wake-up word detection module comprises:
[0043] the second determination sub-module is configured to determine a plurality of posterior probability vectors based on the bone conduction signal and the air conduction signal, the plurality of posterior probability vectors corresponding to a plurality of audio frames included in the bone conduction signal and the air conduction signal in a one-to-one manner, a first posterior probability vector in the plurality of posterior probability vectors being used to indicate a probability that a phoneme of a first audio frame in the plurality of audio frames belongs to a plurality of specified phonemes;
[0044] the wake-up word detection sub-module is configured to detect the wake-up word based on the plurality of posterior probability vectors.
[0045] Optionally, the second determination sub-module is configured to:
[0046] inputting a starting portion of a bone conduction signal and an air conduction signal into the second acoustic model to obtain a first number of bone conduction posterior probability vectors and a second number of air conduction posterior probability vectors output by the second acoustic model, the starting portion of the bone conduction signal being determined according to a detection time delay of the speech detection, the first number of bone conduction posterior probability vectors corresponding to audio frames included in the starting portion of the bone conduction signal in a one-to-one manner, the second number of air conduction posterior probability vectors corresponding to audio frames included in the air conduction signal in a one-to-one manner;
[0047] The first bone conduction posterior probability vector corresponds to a last audio frame of the starting portion of the bone conduction signal, a time length of the last audio frame is less than a frame time length, the first air conduction posterior probability vector corresponds to a first audio frame of the air conduction signal, a time length of the first audio frame is less than the frame time length, and the plurality of posterior probability vectors include the second posterior probability vector, vectors other than the first bone conduction posterior probability vector in the first quantity of bone conduction posterior probability vectors, and vectors other than the first air conduction posterior probability vector in the second quantity of air conduction posterior probability vectors.
[0048] Optionally, the second determining sub-module is configured to:
[0049] The starting portion of the bone conduction signal is input into the third acoustic model according to a detection time delay of the voice detection, and / or
[0050] The bone conduction signal and the air conduction signal are input into the third acoustic model.
[0051] Optionally, the wake-up word detection sub-module is configured to:
[0052] The confidence that the phoneme sequence corresponding to the command word includes the phoneme sequence corresponding to the wake-up word is determined based on the plurality of posterior probability vectors and the phoneme sequence corresponding to the wake-up word.
[0053] In a case where the confidence exceeds a confidence threshold, it is determined that the command word including the wake-up word is detected.
[0054] Optionally, the wake-up word detection sub-module is configured to:
[0055] The confidence that the phoneme sequence corresponding to the command word includes the phoneme sequence corresponding to the wake-up word is determined based on the plurality of posterior probability vectors and the phoneme sequence corresponding to the wake-up word.
[0056] In a case where the confidence exceeds a confidence threshold and the plurality of posterior probability vectors and a plurality of template vectors satisfy a distance condition, it is determined that the command word including the wake-up word is detected, the plurality of template vectors indicating probabilities that phonemes of a voice signal containing complete information of the wake-up word belong to a plurality of specified phonemes.
[0057] Optionally, in a case where the plurality of posterior probability vectors and the plurality of template vectors are in one-to-one correspondence, the distance condition includes that a mean value of distances between the plurality of posterior probability vectors and corresponding template vectors is less than a distance threshold.
[0058] Optionally, the apparatus further includes:
[0059] The acquisition module is configured to acquire a bone-conduction registration signal, the bone-conduction registration signal comprising complete information of the wake-up word.
[0060] The determination module is configured to determine, based on the bone-conduction registration signal and a phoneme sequence corresponding to the wake-up word, a confidence threshold and a plurality of template vectors.
[0061] Optionally, the determination module comprises:
[0062] The third determination submodule is configured to determine, based on the bone-conduction registration signal, a fusion registration signal.
[0063] The fourth determination submodule is configured to determine, based on the fusion registration signal and the phoneme sequence corresponding to the wake-up word, the confidence threshold and the plurality of template vectors.
[0064] Optionally, the fourth determination submodule is configured to:
[0065] input a plurality of registration audio frames comprised in the fusion registration signal into the first acoustic model to obtain a plurality of registration posterior probability vectors output by the first acoustic model, the plurality of registration posterior probability vectors corresponding one-to-one to the plurality of registration audio frames, a first registration posterior probability vector in the plurality of registration posterior probability vectors indicating a probability that a phoneme of a first registration audio frame in the plurality of registration audio frames belongs to a plurality of specified phonemes;
[0066] determine the plurality of registration posterior probability vectors as the plurality of template vectors;
[0067] determine the confidence threshold based on the plurality of registration posterior probability vectors and the phoneme sequence corresponding to the wake-up word.
[0068] Optionally, the apparatus further comprises:
[0069] The acquisition module is configured to acquire an air-conduction registration signal.
[0070] The determination module comprises:
[0071] The fifth determination submodule is configured to determine, based on the bone-conduction registration signal and the air-conduction registration signal, a plurality of registration posterior probability vectors, the plurality of registration posterior probability vectors corresponding one-to-one to a plurality of registration audio frames comprised in the bone-conduction registration signal and the air-conduction registration signal, a first registration posterior probability vector in the plurality of registration posterior probability vectors indicating a probability that a phoneme of a first registration audio frame in the plurality of registration audio frames belongs to a plurality of specified phonemes.
[0072] The sixth determination submodule is configured to determine, based on the plurality of registration posterior probability vectors and the phoneme sequence corresponding to the wake-up word, the confidence threshold.
[0073] In a third aspect, an electronic device is provided, which includes a processor and a memory, the memory being configured to store a program for executing the voice wake-up method provided in the first aspect and store data used in implementing the voice wake-up method provided in the first aspect. The processor is configured to execute the program stored in the memory. The storage device can further include a communication bus for establishing a connection between the processor and the memory.
[0074] In a fourth aspect, a computer readable storage medium is provided, which stores instructions, when executed on a computer, causing the computer to execute the voice wake-up method provided in the first aspect.
[0075] In a fifth aspect, a computer program product is provided, which includes instructions, when executed on a computer, causing the computer to execute the voice wake-up method provided in the first aspect.
[0076] The technical effects obtained by the second aspect, the third aspect, the fourth aspect and the fifth aspect are similar to the technical effects obtained by the corresponding technical means in the first aspect, and will not be repeated here.
[0077] The technical solutions provided in the embodiments of the present application can at least bring the following beneficial effects:
[0078] In the embodiments of the present application, the bone conduction microphone is used to collect bone conduction signals for voice detection, which can ensure low power consumption. In addition, considering that the delay of voice detection can cause the collected air conduction signals to be lost, thereby not containing the complete information of the command word input by the sound source, the bone conduction signals collected by the bone conduction microphone contain the command word information input by the sound source, i.e. the bone conduction signals are not lost. Therefore, the present solution is based on the bone conduction signals to detect the wake-up word. In this way, the recognition accuracy of the wake-up word is high, and the accuracy of the voice wake-up is high. BRIEF DESCRIPTION OF DRAWINGS
[0079] Figure 1 is a structural schematic diagram of an acoustic model provided by the embodiments of the present application;
[0080] Figure 2 is a system architecture diagram related to a voice wake-up method provided by the embodiments of the present application;
[0081] Figure 3 is a structural schematic diagram of an electronic device provided by the embodiments of the present application;
[0082] Figure 4 is a flowchart of a voice wake-up method provided by the embodiments of the present application;
[0083] Figure 5is a principle diagram of bone conduction signal and air conduction signal generation provided by an embodiment of the present application;
[0084] Figure 6 is a signal timing diagram provided by an embodiment of the present application;
[0085] Figure 7 is a method diagram of signal splicing provided by an embodiment of the present application;
[0086] Figure 8 is a diagram of down-sampling of bone conduction signal provided by an embodiment of the present application;
[0087] Figure 9 is a diagram of gain adjustment of bone conduction signal provided by an embodiment of the present application;
[0088] Figure 10 is a method diagram of training a generation network model provided by an embodiment of the present application;
[0089] Figure 11 is a structure diagram of another acoustic model provided by an embodiment of the present application;
[0090] Figure 12 is a structure diagram of another acoustic model provided by an embodiment of the present application;
[0091] Figure 13 is a method flow chart of another voice wake-up provided by an embodiment of the present application;
[0092] Figure 14 is a method flow chart of another voice wake-up provided by an embodiment of the present application;
[0093] Figure 15 is a method flow chart of another voice wake-up provided by an embodiment of the present application;
[0094] Figure 16 is a method flow chart of another voice wake-up provided by an embodiment of the present application;
[0095] Figure 17 is a method flow chart of another voice wake-up provided by an embodiment of the present application;
[0096] Figure 18 is a method flow chart of another voice wake-up provided by an embodiment of the present application;
[0097] Figure 19 is a method flow chart of another voice wake-up provided by an embodiment of the present application;
[0098] Figure 20 is a method flow chart of another voice wake-up provided by an embodiment of the present application;
[0099] Figure 21 is another method flowchart of the method for registering the wake-up word provided by the embodiment of the present application;
[0100] Figure 22 is another method flowchart of the method for registering the wake-up word provided by the embodiment of the present application;
[0101] Figure 23 is another method flowchart of the method for registering the wake-up word provided by the embodiment of the present application;
[0102] Figure 24 is another method flowchart of the method for registering the wake-up word provided by the embodiment of the present application;
[0103] Figure 25 is a method schematic diagram for training the first acoustic model provided by the embodiment of the present application;
[0104] Figure 26 is another method schematic diagram for training the first acoustic model provided by the embodiment of the present application;
[0105] Figure 27 is another method schematic diagram for training the first acoustic model provided by the embodiment of the present application;
[0106] Figure 28 is another method schematic diagram for training the first acoustic model provided by the embodiment of the present application;
[0107] Figure 29 is a method schematic diagram for training the second acoustic model provided by the embodiment of the present application;
[0108] Figure 30 is a method schematic diagram for training the third acoustic model provided by the embodiment of the present application;
[0109] Figure 31 is a structure schematic diagram of the voice wake-up device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0110] In order to make the purpose, technical scheme and advantages of the embodiments of the present application more clear, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0111] In order to facilitate understanding, first, some names or terms in the embodiments of the present application are explained.
[0112] Speech recognition: also known as automatic speech recognition (ASR). Speech recognition refers to recognizing the content of the words contained in the speech signal by a computer.
[0113] Voice wake-up: also known as keyword spotting (KWS), wake-word detection, wake-word recognition, etc. Voice wake-up refers to real-time detection of a wake-up word in a continuous speech stream, and waking up an intelligent device when a named word input by a sound source is the wake-up word.
[0114] Deep learning (DL): a learning algorithm based on data representation in machine learning.
[0115] Next, some related knowledge involved in speech recognition in the embodiments of the present application is introduced.
[0116] Voice activate detector (VAD)
[0117] VAD is used to determine when there is voice input and when it is a silent state, and is also used to cut out the effective segment with voice input. The subsequent operation of speech recognition is performed on the effective segment cut out by VAD, so as to reduce the noise misrecognition rate of the speech recognition system and the system power consumption. In a near-field environment, due to limited attenuation of the voice signal, the signal-noise ratio (SNR) is relatively high, and only a simple method (such as zero-crossing rate, signal energy) is needed to do voice activation detection. However, in a far-field environment, due to the relatively long transmission distance of the voice signal, the attenuation is relatively serious, and thus the SNR of the microphone collected data is very low. In this case, the simple voice activation detection method has poor effect. Using deep neural networks (DNN) for voice activation detection is a commonly used method in deep learning-based speech recognition systems. VAD is an implementation way of voice detection, and the embodiments of the present application take voice detection through VAD as an example for introduction. In other embodiments, voice detection can also be performed through other ways.
[0118] Speech recognition
[0119] For a speech recognition system, the first step is to detect whether there is speech input, i.e., voice activity detection (VAD). In a low-power design, VAD employs an always-on working mechanism compared to other parts of the speech recognition. When VAD detects speech input, VAD will wake up the subsequent recognition system. The recognition system mainly includes feature extraction, recognition modeling, and decoding to obtain the recognition result, etc. Among them, model training includes acoustic model training, language model training, etc. Speech recognition is essentially a process of converting an audio sequence into a text sequence, i.e., given a speech input, find the text sequence with the highest probability. Based on the Bayes principle, the speech recognition problem can be decomposed into the conditional probability of the given text sequence appearing this speech and the prior probability of the text sequence appearing. The model obtained by modeling the conditional probability is the acoustic model, and the model obtained by modeling the prior probability of the text sequence appearing is the language model.
[0120] It should be noted that in order to analyze and recognize the speech signal, the speech signal needs to be framed, that is, the speech signal is cut into multiple small segments, each small segment is called a frame. The framing operation is generally not simply cutting, but using a window function to achieve it. After framing, there is generally an overlap between adjacent frames. In an embodiment of the present application, a frame of audio is obtained by framing, and framing is to enable the acoustic model to analyze the sound signal. For example, the speech signal is framed using a window function, assuming that the window function indicates framing with a frame length of 25 ms (milliseconds) and a frame shift of 10 ms, then after framing, each frame of audio has a length of 25 ms, and there is an overlap of 25-10=15 ms between adjacent two frames.
[0121] Two concepts are explained here. Phonemes: the pronunciation of a word is composed of phonemes, which is a unit of pronunciation. The phoneme set (i.e., pronunciation dictionary) of English, such as a set of 39 phonemes from Carnegie Mellon University. The phoneme set of Chinese, such as directly using all initial consonants and vowels as the phoneme set, or if there are tones and no tones, the phoneme set includes more phonemes. For example, in an embodiment of the present application, the phoneme set includes 100 phonemes. State: can be regarded as a more detailed speech unit than phonemes, usually one phoneme is divided into 3 states. In an embodiment of the present application, one frame of audio corresponds to one phoneme, and a plurality of phonemes form a word (character). Then, as long as it is known which phoneme each frame of audio corresponds to, the result of speech recognition is also obtained. In some embodiments, a plurality of frames of audio correspond to a state, and every three states combine to form a phoneme, and a plurality of phonemes combine to form a word. Then, as long as it is known which state each frame of audio corresponds to, the result of speech recognition is also obtained.
[0122] Acoustic model, decoding, and speech wake-up
[0123] In speech recognition, taking an audio frame corresponding to a phoneme as an example, the acoustic model can know the phoneme corresponding to each audio frame as the probability of each phoneme in the phoneme set, that is, the posterior probability vector corresponding to the audio. In a popular way, a large number of parameters are stored in the acoustic model, and through these parameters, the posterior probability vector corresponding to each audio frame can be known. These parameters can be obtained by training the acoustic model, and the training of the acoustic model requires a large amount of speech data. After obtaining the posterior probability vector corresponding to each audio frame through the acoustic model, a decoding graph (also called state network, search space, etc.) is constructed based on a language model, a pronunciation dictionary, etc., the posterior probability vector corresponding to the continuous multiple frames of audio output by the acoustic model is taken as the input of the decoding graph, and the optimal path in the decoding graph is searched, and the probability of the phoneme corresponding to the speech on the path is the largest. After searching the optimal path, the phoneme corresponding to each audio frame can be known, and the best word string recognized by the speech can be obtained. Among them, the process of searching the optimal path in the state network to obtain the word string can be considered as a kind of decoding, and the decoding is to determine what the word string corresponding to the speech signal is.
[0124] In the decoding of the voice wake-up in the embodiment of the present application, the probability of each phoneme on the decoding path is searched in the decoding graph, and the probabilities of the searched each phoneme are added to obtain a path score. Among them, the decoding path refers to the phoneme sequence corresponding to the wake-up word. If the path score is larger, it is considered that the command word including the wake-up word is detected. That is, the decoding in the embodiment of the present application is to judge whether the word string corresponding to the speech signal is the wake-up word based on the decoding graph.
[0125] To explain the embodiments of the present application, the acoustic model involved in the embodiments of the present application is first further introduced. The acoustic model is a model capable of recognizing a single phoneme, which can be modeled by a hidden Markov model (HMM). The acoustic model is a trained model, which can be trained by using acoustic features of a sound signal and corresponding labels. The acoustic model establishes a corresponding probability distribution between an acoustic signal and a modeling unit, such as an HMM state, a phoneme, a syllable, a word, etc. The modeling unit can also be referred to as a pronunciation unit, and the structure of the acoustic model is, for example, GMM-HMM, DNN-HMM, DNN-CTC, etc. Among them, GMM (gaussian mixed model) represents a Gaussian mixed model, DNN represents a deep neural network, and CTC (connectionist temporal classification) represents a neural network-based temporal classification. In the embodiments of the present application, the modeling unit is taken as a phoneme, and the acoustic model is taken as a DNN-HMM model for example. It should be noted that in the embodiments of the present application, the acoustic model can process each frame of audio, and output a probability that each audio frame belongs to a plurality of specified phonemes, which can be determined according to a pronunciation dictionary. For example, if the pronunciation dictionary includes 100 phonemes, the plurality of specified phonemes are the 100 phonemes.
[0126] Figure 1 FIG. 1 is a structural schematic diagram of an acoustic model provided by the embodiments of the present application. The acoustic model is a DNN-HMM model, the dimension of the input layer of the acoustic model is 3, the dimension of the two hidden layers is 5, and the dimension of the output layer is 3. Among them, the dimension of the input layer represents the feature dimension of the input signal, and the dimension of the output layer represents three state dimensions, each state dimension including probabilities corresponding to a plurality of specified phonemes.
[0127] Then further introduce the decoding. The decoding in speech recognition can be divided into dynamic decoding and static decoding. In the process of dynamic decoding, the language score is dynamically searched in the language model with the dictionary tree as the center. Static decoding refers to that the language model is statically compiled into the decoding graph in advance, and the decoding efficiency is improved through a series of optimization operations such as determination, weight migration and minimization. Exemplarily, the static decoding is adopted in the embodiment of the application, such as the weighted finite state transducer (WFST), and the static decoding based on the HCLG network eliminates redundant information. In the embodiment of the application, the generation of the HCLG network needs to represent the language model, the pronunciation dictionary and the acoustic model into the corresponding FST format, and then compile a large decoding graph through a series of operations such as combination, determination and minimization. The HCLG network construction process is HCLG = ASL(min(RDS(det(H'o min(det(C o min(det(L o G)))))))). Wherein, ASL represents adding a self-loop, min represents minimization, RDS represents removing disambiguation symbols, det represents determination, H' represents the HMM without a self-loop, and o represents combination.
[0128] In the decoding process, the viterbi algorithm is used to find the optimal path in the decoding graph, and there will be no two same paths in the decoding graph. In the decoding process, cumulative beam pruning is adopted, that is, the beam value is subtracted from the current probability maximum path score as a threshold, and the path less than the threshold is pruned. At the same time, the frame synchronous decoding algorithm is adopted to find the starting node of the decoding graph, create the token corresponding to the node, expand the null edge (i.e. the input does not correspond to the real modeling unit) from the token corresponding to the starting node, bind the corresponding token to each reachable node, prune and retain the active token. Every time a frame of audio is input, a token is taken out from the current active token, the corresponding node starts to expand the subsequent non-null edge (i.e. the input corresponds to the real physical modeling unit), and after traversing all active tokens, the current frame active token is pruned and retained. Repeat the above steps until all audio frames are expanded, that is, the token with the maximum score is found, and the final recognition result is obtained by backtracking.
[0129] Network model
[0130] In the embodiments of the present application, the network model refers to the acoustic model described above. The network model used for recognizing the speech signal includes a hidden Markov model (HMM), a Gaussian mixture model (GMM), a deep neural network (DNN), a deep belief networks-HMM (DBN-HMM), a recurrent neural network (RNN), a long short-term memory (LSTM) network, a convolutional neural network (CNN), and the like. In the embodiments of the present application, the CNN and the HMM are used.
[0131] The hidden Markov model is a statistical model, and is currently mainly used in the field of speech signal processing. In the model, whether a state in the Markov chain is transferred to another state depends on the state transition probability, and an observation value generated by a certain state depends on the state generation probability. When performing speech recognition, the HMM first establishes a sound production model for each recognition unit, obtains a state transition probability matrix and an output probability matrix through long-time training, and makes a decision according to the maximum probability in the state transition process during recognition.
[0132] The basic structure of the convolutional neural network includes two parts. One part is a feature extraction layer, and the input of each neuron is connected to a local receptive field of the previous layer, and the feature of the local receptive field is extracted. The other part is a feature mapping layer, and each calculation layer of the network is composed of multiple feature mappings, and each feature mapping is a plane, and the weights of all neurons on the plane are equal. The feature mapping structure uses a function with a small influence function kernel (such as a sigmoid function) as the activation function of the convolutional network, so that the feature mapping has the shift invariance. In addition, since the neurons on a mapping plane share the weights, the number of free parameters of the network is reduced. Each convolutional layer in the convolutional neural network can be followed by a calculation layer for local averaging and secondary extraction. The unique two-time feature extraction structure reduces the feature resolution.
[0133] The loss function is the iterative basis of the network model in training. The loss function is used to evaluate the degree of difference between the predicted value of the network model and the true value. The selection of the loss function affects the performance of the network model. Different network models generally use different loss functions. The loss function can be divided into an empirical risk loss function and a structural risk loss function. The empirical risk loss function refers to the difference between the prediction result and the actual result. The structural risk loss function refers to the empirical risk loss function plus a regularization term. The cross-entropy loss function (cross-entropy loss function) is used in the embodiments of the present application, that is, the CE loss function. The cross-entropy loss function is also a kind of log-likelihood function in nature, which can be used in binary classification and multi-classification tasks. When using sigmoid as the activation function, the cross-entropy loss function is commonly used instead of the mean square error loss function, because the cross-entropy loss function can perfectly solve the problem of slow weight update of the square loss function, and has the good property of fast weight update when the error is large and slow weight update when the error is small.
[0134] The network model propagates the error in the reverse direction, uses the loss function, and adopts the gradient descent method to adjust the network parameters. The gradient descent method is an optimization algorithm, and the central idea is to update the parameter value in the direction of the gradient of the objective function to hope to reach the minimum (or maximum) of the objective function. The gradient descent method is a commonly used optimization algorithm in deep learning. Gradient descent is for the loss function, and the purpose is to find the minimum value of the loss function as soon as possible. The core of the back propagation algorithm is to define a special variable of neuron error. The neuron error is propagated layer by layer from the output layer. Then the partial derivative of the weight and bias is calculated by the formula using the neuron error. Gradient descent is a way to solve the minimum value problem, and back propagation is a way to solve the gradient calculation.
[0135] Figure 2 is a system architecture diagram involved in a voice wake-up method provided by the embodiments of the present application. Referring to Figure 2 , the system architecture includes a wearable device 201 and a smart device 202. The wearable device 201 and the smart device 202 are connected through wired or wireless means to communicate. Among them, the smart device 202 is the device to be woken up in the embodiments of the present application.
[0136] In the embodiments of the present application, the wearable device 201 is configured to receive a voice signal, and send an instruction to the smart device 202 based on the received voice signal. The smart device 202 is configured to receive the instruction sent by the wearable device 201, and perform a corresponding operation based on the received instruction. For example, the wearable device 201 is configured to collect a voice signal, detect a command word contained in the collected voice signal, and if the command word is detected to include a wake-up word, send a wake-up instruction to the smart device 202 to wake up the smart device 202. The smart device 202 is configured to receive the wake-up instruction, and then enter a working state from a sleep state.
[0137] In the embodiments of the present application, the wearable device 201 is configured to receive a voice signal, and send an instruction to the smart device 202 based on the received voice signal. The smart device 202 is configured to receive the instruction sent by the wearable device 201, and perform a corresponding operation based on the received instruction. For example, the wearable device 201 is configured to collect a voice signal, detect a command word contained in the collected voice signal, and if the command word is detected to include a wake-up word, send a wake-up instruction to the smart device 202 to wake up the smart device 202. The smart device 202 is configured to receive the wake-up instruction, and then enter a working state from a sleep state.
[0138] In the embodiments of the present application, the wearable device 201 is, for example, a wireless earphone, smart glasses, a smart watch, a smart bracelet, or the like. The smart device 202 (i.e., a device to be woken up) is, for example, a smart speaker, a smart home appliance, a smart toy, a smart robot, or the like. Optionally, in some embodiments, the wearable device 201 and the smart device 202 are the same device.
[0139] It should be noted that the system architecture and business scenarios described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. It can be known by those skilled in the art that, with the evolution of system architecture and the appearance of new business scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0140] For reference Figure 3 , Figure 3 is a structural schematic diagram of an electronic device according to an embodiment of the present application. Optionally, the electronic device is the wearable device 201 shown in Figure 2 . The electronic device includes one or more processors 301, a communication bus 302, a memory 303, one or more communication interfaces 304, a bone conduction microphone 308, and an air microphone 309.
[0141] The processor 301 is a general central processing unit (CPU), a network processing (NP), a microprocessor, or one or more integrated circuits for implementing the schemes of the present application, such as an application-specific integrated circuit (ASIC), a programmable logic device (PLD) or a combination thereof. Alternatively, the PLD is a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.
[0142] The communication bus 302 is used to transmit information between the above components. Alternatively, the communication bus 302 is divided into an address bus, a data bus, a control bus and the like. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0143] Alternatively, the memory 303 is a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), an optical disk (including a compact disc read-only memory (CD-ROM), a compressed disk, a laser disk, a digital versatile disk, a Blu-ray disk and the like), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but is not limited thereto. The memory 303 exists independently and is connected to the processor 301 through the communication bus 302, or the memory 303 is integrated with the processor 301.
[0144] The communication interface 304 uses any transceiver-like mechanism for communicating with other devices or communication networks. The communication interface 304 includes a wired communication interface, and optionally also includes a wireless communication interface. The wired communication interface is, for example, an Ethernet interface, etc. Optionally, the Ethernet interface is an optical interface, an electrical interface, or a combination thereof. The wireless communication interface is, for example, a wireless local area networks (WLAN) interface, a cellular network communication interface, or a combination thereof, etc.
[0145] Optionally, in some embodiments, the electronic device includes multiple processors, such as the processor 301 and the processor 305 as shown in FIG. 3. Each of the processors is a single-core processor, or a multi-core processor. Optionally, the processor herein refers to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). Figure 2
[0146] In some embodiments, the electronic device also includes an output device 306 and an input device 307. The output device 306 is in communication with the processor 301 and is capable of displaying information in various manners. For example, the output device 306 is a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device 307 is in communication with the processor 301 and is capable of receiving user input in various manners. For example, the input device 307 includes one or more of a mouse, a keyboard, a touch screen device, or a sensor device, etc.
[0147] In the embodiments of the present application, the input device 307 includes a bone conduction microphone 308 and an air microphone 309, which are respectively used to collect bone conduction signals and air conduction signals. The processor 301 is configured to wake up the smart device by the voice wake-up method provided in the embodiments of the present application based on the bone conduction signals or based on the bone conduction signals and the air conduction signals. Optionally, after waking up the smart device, the processor 301 is further configured to control the smart device to perform a task based on the bone conduction signals, or the air conduction signals, or the bone conduction signals and the air conduction signals.
[0148] In some embodiments, the memory 303 is configured to store program code 310 for implementing the solutions of the present application, and the processor 301 is capable of executing the program code 310 stored in the memory 303. The program code 310 includes one or more software modules, and the electronic device is capable of implementing the voice wake-up method provided in the embodiments of the present application by means of the processor 301 and the program code 310 in the memory 303. Figure 4
[0149] Figure 4 is a flowchart of a voice wake-up method provided by an embodiment of the present application, which is applied to a wearable device. Please refer to Figure 4 , which comprises the following steps.
[0150] Step 401: Voice detection is performed according to a bone conduction signal collected by a bone conduction microphone, the bone conduction signal containing command word information input by a sound source.
[0151] As known from the foregoing, in order to realize low power consumption of the wearable device, before the smart device (i.e. the device to be woken up) is woken up, since the bone conduction microphone has low power consumption, the bone conduction signal can be collected by the bone conduction microphone, and voice detection (such as voice activity detection VAD) is performed based on the bone conduction signal, so as to detect whether there is voice input. In the case where no voice input is detected, the components in the wearable device other than the bone conduction microphone can be in a sleep state, so as to reduce power consumption, and in the case where voice input is detected, the other components of the wearable device are controlled to be turned on. For example, in the case where the wearable device is also provided with an air microphone, since the air microphone is a device with high power consumption, in order to reduce power consumption, the air microphone is controlled to be turned on and turned off, and when voice input is detected (such as when a user is speaking), the air microphone is turned on to perform sound pickup (i.e. to collect an air conduction signal), so as to reduce power consumption of the wearable device. That is, before the smart device is woken up, the air microphone is in a sleep state to reduce power consumption, and in the case where voice input is detected, the air microphone is turned on.
[0152] Among them, there can be multiple implementation manners of voice activity detection performed by the wearable device based on the bone conduction signal collected by the bone conduction microphone, which are not limited by the embodiments of the present application. Next, some implementation manners of voice activity detection are exemplarily introduced. It should be noted that voice activity detection is mainly used to detect whether there is human voice signal in the current input signal. Among them, voice activity detection judges the input signal, so as to distinguish voice segments from non-voice segments (such as segments containing only various background noise signals), so that different processing methods can be taken for different segments of signals.
[0153] Optionally, the voice activation detection detects whether there is voice input by extracting features of the input signal. For example, whether there is voice input is detected by extracting features of short time energy (STE) and short time zero cross counter (ZCC) of each frame of the input signal, i.e., voice activation detection is performed based on energy-based features. The short time energy refers to the energy of a frame of signal, and the zero cross rate refers to the number of times that a frame of time domain signal crosses 0 (time axis). For another example, some high-precision VADs extract multiple features such as energy-based features, frequency domain features, cepstrum features, harmonic features, and long-time features for comprehensive detection. Optionally, in addition to extracting features, a threshold comparison or a statistical method or a machine learning method can be combined to determine whether a frame of input signal is a voice signal or a non-voice signal. Next, the energy-based features, the frequency domain features, the cepstrum features, the harmonic features, and the long-time features are introduced respectively.
[0154] Energy-based features: VAD is performed based on two features of STE and ZCC. In the case of a large signal-to-noise ratio (SNR), the STE of a voice segment is relatively large and the ZCC is relatively small, and the STE of a non-voice segment is relatively small and the ZCC is relatively large. Because a voice signal with human voice usually has large energy, and most of the energy is contained in the low frequency band, and a noise signal usually has small energy and contains more high frequency information. Therefore, the two features of the input signal can be extracted to distinguish voice signals and non-voice signals. The method for calculating the STE can be to calculate the sum of squares of the energy of each frame of the input signal through a spectrum diagram. The method for calculating the short time zero cross rate can be to calculate the number of zero crossings corresponding to each frame of the input signal in the time domain. For example, all sampling points in the frame are shifted one point to the left or right in the time domain, the amplitude values of the shifted sampling points and the amplitude values of the unshifted sampling points are multiplied at the corresponding points, and if the sign of the product of the corresponding two sampling points is negative, it means that the corresponding sampling point crosses zero. The number of products that are negative in the frame is calculated to obtain the short time zero cross rate.
[0155] Frequency domain features: The time domain signal of the input signal is converted into a frequency domain signal by a short time Fourier transform or other time-frequency transform method to obtain a spectrum diagram, and the frequency domain features are obtained based on the spectrum diagram. For example, the envelope features of the frequency band are extracted based on the spectrum diagram. In some experiments, when the SNR is 0 dB, the long-time envelope of some frequency bands can distinguish voice segments and noise segments.
[0156] Cepstral features: such as including energy cepstral peak. For VAD, the energy cepstral peak determines the pitch of the speech signal. In some embodiments, Mel-frequency cepstral coefficients (MFCC) are used as the cepstral features.
[0157] Harmonic-based features: one of the prominent features of speech signal is that it contains the fundamental frequency and its multiple harmonic frequencies, even in strong noise scenarios, this feature is present. The autocorrelation method can be used to find the fundamental frequency of the speech signal.
[0158] Long-term features: speech signal is a non-stationary signal, and the normal speech rate usually emits 10-15 phonemes per second, and the spectral distribution between phonemes is different, which leads to the change of speech statistical characteristics over time. However, most of the noise in daily life is stationary, that is, it changes slowly, such as white noise. Based on this, long-term features can be extracted to determine whether the input signal is a speech signal or a non-speech signal.
[0159] It should be noted that in the embodiments of the present application, the input signal for voice activation detection is a bone conduction signal collected by a bone conduction microphone. Voice activation detection is performed on each frame of the received bone conduction signal to detect whether there is a voice input. Since the bone conduction microphone is always in working state, the bone conduction signal continuously collected by the bone conduction microphone contains complete information of the command word input by the sound source, that is, the bone conduction signal does not lose the head.
[0160] Optionally, the sampling rate of the bone conduction signal is 32 kHz (kilohertz), 48 kHz, etc., which is not limited in the embodiments of the present application. The sensor in the bone conduction microphone is a non-acoustic sensor, which can shield the influence of ambient noise and has strong anti-noise performance.
[0161] Step 402: In the case of detecting a voice input, the wake-up word is detected based on the bone conduction signal.
[0162] In the embodiments of the present application, in the case of detecting a voice input, the wearable device detects the wake-up word based on the bone conduction signal to detect whether the command word includes a wake-up word. It should be noted that there are many implementation manners for the wearable device to detect the wake-up word based on the bone conduction signal, and two implementation manners are introduced as follows.
[0163] First implementation manner
[0164] In the embodiments of the present application, the implementation manner of the wearable device to detect the wake-up word based on the bone conduction signal is to determine a fusion signal based on the bone conduction signal, and to detect the wake-up word based on the fusion signal.
[0165] Firstly, an implementation manner of determining a fusion signal based on a bone conduction signal by a wearable device is introduced. It should be noted that there are multiple manners of determining the fusion signal based on the bone conduction signal by the wearable device, and four of the manners are introduced next.
[0166] Way 1 of determining a fusion signal based on a bone conduction signal Before determining the fusion signal based on the bone conduction signal, an air microphone is opened, and an air conduction signal is collected by the air microphone. For example, in a case where it is detected that there is voice input, the air microphone is opened, and the air conduction signal is collected by the air microphone. The wearable device fuses a starting part of the bone conduction signal and the air conduction signal to obtain the fusion signal. The starting part of the bone conduction signal is determined according to a detection time delay of voice detection (such as VAD). That is, the wearable device collects the bone conduction signal and the air conduction signal, and performs head loss compensation on the air conduction signal by using the starting part of the bone conduction signal, so that the obtained fusion signal also contains command word information of a sound source input. In addition, the length of the fusion signal is relatively short, which can reduce the data processing amount to a certain extent. Alternatively, in the embodiment of the present application, signal fusion is performed by signal splicing, and in some embodiments, signal fusion can also be performed by signal superposition and the like. In the following embodiments, signal fusion by signal splicing is taken as an example for introduction.
[0167] It should be noted that the bone conduction signal and the air conduction signal are signals generated by the same sound source, and the transmission paths of the bone conduction signal and the air conduction signal are different. As shown in FIG. 3, Figure 5 the bone conduction signal is a signal formed by a vibration signal (excitation signal) being transmitted through the internal bones and tissues of the human body, and the air conduction signal is a signal formed by a sound wave being transmitted through the air.
[0168] Figure 6 is a signal timing diagram provided by an embodiment of the present application. The signal timing diagram shows the timing relationship of the bone conduction signal, the air conduction signal, the VAD control signal and the user voice signal. When the sound source emits a voice signal, the bone conduction signal immediately becomes a high-level signal, and after a time of △t, the VAD determines that there is voice input, at this time, the VAD control signal is generated, the VAD control signal controls the air microphone to be opened, and the air conduction signal is collected, that is, at this time, the air conduction signal becomes a high-level signal. It can be seen that the bone conduction signal and the user voice signal change synchronously, and the air conduction signal has a delay of △t compared with the bone conduction signal, and the delay is caused by the detection time delay of the VAD. Wherein, △t represents the detection time delay of the voice activation detection, that is, the time difference between the time when it is detected that there is voice input and the time when the user actually inputs the voice.
[0169] It should be noted that in the embodiments of the present application, the VAD can detect the speech segments and non-speech segments in the bone conduction signal, and the endpoint detection can detect the speech segments and non-speech segments in the air conduction signal. Before the wearable device fuses the starting part of the bone conduction signal and the air conduction signal, the wearable device cuts out the speech segments in the bone conduction signal based on the detection result of the VAD, cuts out the speech segments in the air conduction signal based on the detection result of the endpoint detection, fuses the starting part of the cut-out speech segments in the bone conduction signal with the cut-out speech segments in the air conduction signal, to obtain a fused signal. For example, the time range of the cut-out speech segments in the bone conduction signal is [0, t], the time range of the starting part of the bone conduction signal (i.e., the starting part of the cut-out speech segments) is [0, △t], the time range of the cut-out speech segments in the air conduction signal is [△t, t], and the time length of the obtained fused signal is t. Wherein, △t represents the detection delay of the voice activity detection, and t represents the total time length of the actual speech input. Figure 5
[0170] Figure 7 is a schematic diagram of a signal fusion method provided by the embodiments of the present application. For example, the signal fusion is performed by signal splicing, referring to Figure 7 , x1[n] represents the starting part of the bone conduction signal, x2[n] represents the air conduction signal, f(x) represents a splicing function, and f(x): b[n] 0,t = concat[x1[n] 0,△t +x2[n] 0,△t , x2[n] △t,t ], wherein x2[n] 0,△t is zero. That is, the starting part of the bone conduction signal (i.e., the speech segments from 0 to △t) is spliced with the air conduction signal (i.e., the speech segments from △t to t) by f(x) to obtain the fused signal b[n].
[0171] Optionally, before the initial part of the bone conduction signal and the air conduction signal are fused, the wearable device pre-processes the air conduction signal, and the pre-processing includes front-end enhancement. The front-end enhancement can eliminate the effects of some noise and different sound sources, etc., so that the air conduction signal after front-end enhancement can better reflect the essential characteristics of the speech, so as to improve the accuracy of speech wake-up. It should be noted that there are many methods for front-end enhancement of the air conduction signal, for example, endpoint detection and speech enhancement, such as echo cancellation, beamforming algorithm, noise cancellation, automatic gain control, dereverberation, etc. Among them, endpoint detection can distinguish speech segments and non-speech segments of the air conduction signal, that is, accurately determine the starting point of the speech segment. After endpoint detection, only the speech segment of the air conduction signal can be processed subsequently, which can improve the accuracy and recall rate of speech recognition. Speech enhancement is to eliminate the influence of environmental noise on the speech segment. For example, echo cancellation is to use an effective echo cancellation algorithm to suppress the interference of the far-end signal, mainly including double-talk detection and delay estimation, such as judging the current speaking mode (such as near-talking mode, far-talking mode, double-talking mode, etc.), adjusting the filter based on the current speaking mode Corresponding strategy, and then filtering out the far-end interference in the air conduction signal through the filter, and then eliminating the interference of residual noise through a post-filtering algorithm. For another example, the automatic gain algorithm is used to quickly gain the signal to a suitable volume, and the present scheme can multiply all sampling points of the air conduction signal by a corresponding gain factor through hard gain processing, and multiply a corresponding gain factor at each frequency in the frequency domain. Among them, the frequencies of the air conduction signal can be weighted according to the equal loudness curve, and the loudness gain factor is mapped to the equal loudness curve to determine the gain factor of each frequency.
[0172] Optionally, before fusing the initial part of the bone conduction signal and the air conduction signal, the wearable device pre-processes the bone conduction signal, the pre-processing including downsampling and / or gain adjustment. The downsampling can reduce the data volume of the bone conduction signal and improve the efficiency of data processing, and the gain adjustment is used to enhance the energy of the adjusted bone conduction signal, for example, the gain adjustment makes the average energy of the bone conduction signal consistent with the average energy of the air conduction signal. It should be noted that there are many methods for downsampling and / or gain adjustment of the bone conduction signal, and the embodiments of the present application do not limit them. Downsampling refers to reducing the sampling frequency (also known as sampling rate) of the signal, which is a way of signal resampling. The sampling frequency refers to the number of samples of the sound amplitude extracted per second after the analog sound waveform is digitized. In the process of downsampling the air conduction signal x[n] with a sampling frequency of Fs and including N sampling points, every M-1 sampling points are extracted to obtain an air conduction signal y[m] including M sampling points. According to the Nyquist sampling theorem, downsampling may cause spectral confusion of the signal, so a low-pass de-aliasing filter can be used to process the air conduction signal before downsampling, i.e., anti-aliasing filtering, to reduce the spectral confusion caused by subsequent downsampling. Gain adjustment refers to adjusting the amplitude value of the sampling point of the bone conduction signal through a gain factor, or adjusting the energy value of the frequency point of the bone conduction signal. The gain factor can be determined according to a gain function, or according to the statistical information of the air conduction signal and the bone conduction signal, and the embodiments of the present application do not limit them.
[0173] Figure 8 is a schematic diagram of the present application embodiment for downsampling the bone conduction signal. Referring to Figure 8 , it is assumed that the sampling rate of the bone conduction signal is 48 kHz, and the collected bone conduction signal x[n] is first sent to an anti-aliasing filter H(z) to prevent signal aliasing. v[n] represents the bone conduction signal after the anti-aliasing filter, and the sampling rate remains unchanged. The three times down-sampling is performed on v[n] to obtain the three times down-sampled bone conduction signal y[m], and the sampling rate is reduced to 16 kHz.
[0174] Figure 9 is a schematic diagram of the present application embodiment for gain adjustment of the bone conduction signal. Referring to Figure 9 , x[n] represents the bone conduction signal, f(g) represents the gain function, f(g): y[n] = G*x[n], i.e., the gain adjustment of x[n] is performed by the gain factor G determined by f(g) to obtain the gain-adjusted bone conduction signal y[n].
[0175] Way 2 of determining a fusion signal based on a bone conduction signalBefore determining the fusion signal based on the bone conduction signal, the air microphone is turned on, and the air conduction signal is collected through the air microphone. The wearable device generates an enhanced start signal based on the start part of the bone conduction signal, and fuses the enhanced start signal and the air conduction signal to obtain the fusion signal. The start part of the bone conduction signal is determined according to the detection delay of the voice detection. That is, the wearable device generates an enhanced start signal using the start part of the bone conduction signal, and performs head loss compensation on the collected air conduction signal using the enhanced start signal, so that the obtained fusion signal also contains the command word information of the sound source input. In addition, the length of the fusion signal is relatively short, which can reduce the data processing amount to a certain extent.
[0176] It should be noted that the difference between the above-mentioned manner 1 of determining the fusion signal based on the bone conduction signal and the above-mentioned manner 2 of determining the fusion signal based on the bone conduction signal is that in the manner 2, the enhanced start signal is generated using the start part of the bone conduction signal, and the enhanced start signal is fused with the air conduction signal, rather than fusing the start part of the bone conduction signal with the air conduction signal. In addition to the above, the other contents introduced in the above-mentioned manner 1 are applicable to the above-mentioned manner 2, and will not be described in detail in the manner 2. For example, in the above-mentioned manner 2, the voice segment detection can also be performed on the bone conduction signal and the air conduction signal to intercept the voice segment, and the signal splicing is performed based on the intercepted voice segment, so as to reduce the data processing amount. The wearable device can also preprocess the bone conduction signal and the air conduction signal, such as downsampling and / or gain adjustment of the bone conduction signal, and voice enhancement of the air conduction signal.
[0177] In the embodiment of the present application, the wearable device can input the start part of the bone conduction signal into a generation network model to obtain an enhanced start signal output by the generation network model. The generation network model is a model trained based on a deep learning algorithm, and can be regarded as a kind of signal generator, which can generate a voice signal containing the information of the input signal and close to the real voice based on the input signal. In the embodiment of the present application, the enhanced start signal contains the signal information of the start part of the bone conduction signal, and the enhanced start signal is close to the real voice signal. It should be noted that the network structure, training method, training device, etc. of the generation network model are not limited in the embodiment of the present application. Next, a training method of a generation network model is exemplarily introduced.
[0178] In the embodiments of the present application, taking the generation network model trained on a computer device as an example, the computer device obtains a first training data set, and the first training data set includes a plurality of first sample signal pairs. The computer device inputs the starting part of the bone conduction sample signal in the plurality of first sample signal pairs into an initial generation network model to obtain a plurality of enhanced starting sample signals output by the initial generation network model. The computer device inputs the plurality of enhanced starting sample signals and the starting part of the air conduction sample signal in the plurality of first sample signal pairs into an initial decision network model to obtain a decision result output by the initial decision network model. The computer device adjusts the network parameters of the initial generation network model based on the decision result to obtain a trained generation network model. Wherein, one first sample signal pair includes the starting part of one bone conduction sample signal and the starting part of one air conduction sample signal, one first sample signal pair corresponds to one command word, and the bone conduction sample signal and the air conduction sample signal contain complete information of the corresponding command word.
[0179] Optionally, the first sample signal pair obtained by the computer device contains a bone conduction sample signal and an air conduction sample signal, and the computer device intercepts the starting part of the bone conduction sample signal and the starting part of the air conduction sample signal to obtain the input data of the initial generation network model and the initial decision network model. That is, the computer device first obtains a complete voice signal, and then intercepts the starting part to obtain the training data. Alternatively, the first sample signal pair obtained by the computer device only contains the starting part of the bone conduction sample signal and the starting part of the air conduction sample signal.
[0180] Optionally, the first training data set includes directly collected voice data, public voice data and / or voice data purchased from a third party. Optionally, before training, the computer device can pre-process the obtained first training data set pairs to obtain a pre-processed first training data set, and the pre-processed first training data set can simulate the distribution of real voice data, so as to be closer to the voice in the real scene and increase the diversity of the training samples. Illustratively, the first training data set is backed up, that is, an additional data is added, and the backup data is pre-processed. Optionally, the backup data is divided into multiple parts, and each part of data is pre-processed in one way. The pre-processing of each part of data can be different, which can double the total training data and ensure the comprehensiveness of the data, balance the performance and training cost, and improve the accuracy and robustness of voice recognition to a certain extent. Wherein, the method of pre-processing each part of data can include one or more of noise addition, volume enhancement, reverb addition, time shifting, pitch shifting, time stretching, etc.
[0181] Exemplarily, the increasing noise refers to mixing one or more background noises into the speech signal, so that the training data can cover more kinds of noises, such as office environment noise, canteen environment noise, street environment noise, and the like. Different signal-to-noise ratios can also be mixed in, for example, the signal-to-noise ratio can be selected in a normal distribution manner, so that the mean value of the signal-to-noise ratio is optimal, and the mean value can be 10 dB, 20 dB, or the like, and the signal-to-noise ratio can be from 10 dB to 30 dB, or the like. Wherein, the computer device can calculate the noise energy N based on the signal energy S and the signal-to-noise ratio SNR through the formula SNR = 10*log 10 (S 2 / N 2 ) to calculate the noise energy N. The volume enhancement refers to enhancing or weakening the volume of the speech signal according to the volume variation coefficient, and the value range of the volume variation coefficient can be 0.5 to 1.5, or other value ranges. The increasing reverberation refers to adding reverberation processing to the speech signal, and the reverberation is caused by the reflection of the space environment to the sound signal. The tone change, such as the high pitch correction, is to change the pitch of the speech without affecting the sound speed. The time stretching refers to changing the speed or duration of the speech signal without affecting the pitch, that is, changing the speech speed, so that the training data can cover different speech speeds, and the variation range of the speech speed can be between 0.9 to 1.1 or in other ranges.
[0182] Figure 10 is a schematic diagram of a method for training a generated network model provided by an embodiment of the present application. The generator (i.e., the initial generated network model) is a network for generating a speech signal, and the starting part of the bone conduction sample signal in the first training data set is input into the generator. Optionally, a random noise is superimposed in the bone conduction sample signal before being input into the generator. The input bone conduction sample signal is processed by the generator to generate an enhanced starting sample signal. The discriminator (i.e., the initial decision network model) is a decision network for judging whether the input signal is a real speech signal, and the decision result output by the discriminator indicates whether the input signal is a real speech signal. If the output decision result is 1, it indicates that the discriminator determines that the input signal is a real speech signal, and if the output decision result is 0, it indicates that the discriminator determines that the input signal is not a real speech signal. The parameters in the generator and the discriminator are adjusted to train the generator and the discriminator by judging whether the decision result is accurate. In the training process, the goal of the generator is to generate a fake speech signal to deceive the discriminator, and the goal of the discriminator is to distinguish whether the input signal is real or generated. It can be seen that the generator and the discriminator are essentially playing a game through the training data, and the capabilities of the generator and the discriminator are improved in the process of the game. In an ideal case, the accuracy of the trained discriminator is close to 0.5.
[0183] After the training is completed, the generated network model obtained by the training is deployed to the wearable device, and the wearable device inputs the starting signal of the collected bone conduction signal into the generated network model to obtain an enhanced starting signal output by the generated network model. It should be noted that in addition to the above-mentioned method of generating an enhanced starting signal, the computer device can also generate an enhanced starting signal based on the starting signal of the bone conduction signal by using other methods, and the embodiments of the present application do not make any limitation in this regard.
[0184] Way 3 of determining a fusion signal based on a bone conduction signal Before the fusion signal is determined based on the bone conduction signal, the air microphone is turned on, and the air conduction signal is collected by the air microphone. The wearable device directly fuses the bone conduction signal and the air conduction signal to obtain the fusion signal. In this way, the obtained fusion signal also contains the command word information input by the sound source, and in addition, the fusion signal contains not only the complete speech information in the bone conduction signal, but also the complete speech information in the air conduction signal, so that the speech features contained in the fusion signal are more abundant, and the accuracy of speech recognition is improved to a certain extent.
[0185] It should be noted that the difference between the above-mentioned method 1 of determining the fusion signal based on the bone conduction signal and the method 3 of determining the fusion signal based on the bone conduction signal is that in the method 3 of determining the fusion signal based on the bone conduction signal, the wearable device directly fuses the bone conduction signal and the air conduction signal, and in addition, the other contents introduced in the above-mentioned method 1 are applicable to the method 3, and in the method 3, they will not be introduced in detail one by one. For example, in the method 3, the speech segment detection can also be performed on the bone conduction signal and the air conduction signal to intercept the speech segment, and the intercepted speech segment is fused to reduce the data processing amount. The bone conduction signal and the air conduction signal can also be preprocessed, for example, the bone conduction signal is down-sampled and / or gain-adjusted, and the air conduction signal is endpoint detected and speech enhanced.
[0186] Exemplarily, it is assumed that the signal fusion is performed by signal splicing, x1[n] represents the bone conduction signal, x2[n] represents the air conduction signal, and f(x) represents the splicing function. It is assumed that f(x): b[n] 0,2t = concat[x1[n] 0,t , x2[n] 0,t ], wherein x2[n] 0,△t is zero. That is, the bone conduction signal (speech segment from 0 to t) and the air conduction signal (signal segment from 0 to t) are spliced by f(x) to obtain the fusion signal b[n]. Alternatively, f(x): b[n] 0,2t-△t = concat[x1[n] 0-t , x2[n] △t-t ]. That is, the bone conduction signal (speech segment from 0 to t) and the air conduction signal (signal segment from Δt to t) are spliced by f(x) to obtain the fusion signal b[n].
[0187] Way 4 of determining a fusion signal based on a bone conduction signal The wearable device determines the bone conduction signal as the fusion signal. That is, the detection of the wake-up word can also be performed only by using the bone conduction signal.
[0188] It should be noted that the difference between the above-described manner 1 of determining the fusion signal based on the bone conduction signal is that, in the manner 4 of determining the fusion signal based on the bone conduction signal, the bone conduction signal is directly used as the fusion signal, and the other contents described in the above-described manner 1 are applicable to the manner 4, and the other contents described in the above-described manner 1 are not repeatedly described in detail in the manner 4. For example, in the manner 4, the detection of the speech segment can also be performed on the bone conduction signal, so as to intercept the speech segment, and the intercepted speech segment is used as the fusion signal, so as to reduce the data processing amount. The bone conduction signal can also be preprocessed, for example, down-sampling and / or gain adjustment are performed on the bone conduction signal.
[0189] Next, the implementation manner of the wearable device for recognizing the fusion signal to perform the detection of the wake-up word is described.
[0190] In the embodiment of the present application, the wearable device inputs the plurality of audio frames included in the fusion signal into the first acoustic model, so as to obtain a plurality of posterior probability vectors output by the first acoustic model. The wearable device performs the detection of the wake-up word based on the plurality of posterior probability vectors. The plurality of posterior probability vectors correspond to the plurality of audio frames included in the fusion signal one by one, that is, one posterior probability vector corresponds to one audio frame included in the fusion signal, and a first posterior probability vector in the plurality of posterior probability vectors is used to indicate the probability that the phoneme in a first audio frame in the plurality of audio frames belongs to a plurality of specified phonemes, that is, one posterior probability vector indicates the probability that the phoneme in the corresponding one audio frame belongs to the plurality of specified phonemes. That is, the wearable device processes the fusion signal by using the first acoustic model, so as to obtain the information of the phonemes included in the fusion signal, and then performs the detection of the wake-up word based on the information of the phonemes. Optionally, in the embodiment of the present application, the first acoustic model can be the network model as described above, or a model with other structures. After the wearable device inputs the fusion signal into the first acoustic model, the wearable device obtains the posterior probability vectors corresponding to the respective audio frames included in the fusion signal by processing the respective audio frames included in the fusion signal by using the first acoustic model.
[0191] In the embodiments of the present application, after the wearable device obtains the plurality of posterior probability vectors output by the first acoustic model, the wearable device determines a confidence that the phoneme sequence corresponding to the command word input by the sound source includes the phoneme sequence corresponding to the wake-up word based on the plurality of posterior probability vectors and the phoneme sequence corresponding to the wake-up word. In the case where the confidence exceeds a confidence threshold, it is determined that the command word including the wake-up word is detected. That is, the wearable device decodes the plurality of posterior probability vectors to determine a confidence. The phoneme sequence corresponding to the wake-up word is referred to as a decoding path, the determined confidence is referred to as a path score, and the confidence threshold is referred to as a wake-up threshold.
[0192] Exemplarily, in the embodiments of the present application, after the posterior probability vector corresponding to each audio frame is obtained by the first acoustic model, the plurality of posterior probability vectors corresponding to the plurality of continuous audio frames are input into a decoding graph (also referred to as a state network) constructed based on a language model and a pronunciation dictionary, and the probabilities of each phoneme on the decoding path are found in the decoding graph. The probabilities of each phoneme found are added to obtain a confidence. The decoding path refers to the phoneme sequence corresponding to the wake-up word. If the confidence is greater than a confidence threshold, it is determined that the command word including the wake-up word is detected.
[0193] Optionally, in order to reduce the false wake-up rate, in the case where the confidence exceeds the confidence threshold and the distance condition is met between the plurality of posterior probability vectors and a plurality of template vectors, the wearable device determines that the command word input by the sound source includes the wake-up word. The plurality of template vectors indicate the probabilities that the phonemes of the speech signal containing complete information of the wake-up word belong to a plurality of specified phonemes. That is, the current input speech needs to not only meet the confidence condition, but also match the template. In the embodiments of the present application, the confidence threshold can be pre-set, for example, based on experience or determined according to the bone conduction registration signal and / or the air conduction registration signal containing complete information of the wake-up word when the wake-up word is registered. The specific implementation manner is introduced below. The plurality of template vectors are the registration posterior probability vectors determined according to the bone conduction registration signal and / or the air conduction registration signal. The specific implementation manner is introduced below.
[0194] Optionally, in a case that the plurality of posterior probability vectors correspond to the plurality of template vectors one-to-one, the distance condition comprises: a mean value of distances between the plurality of posterior probability vectors and corresponding template vectors is less than a distance threshold. It should be noted that, if the plurality of posterior probability vectors correspond to the plurality of template vectors one-to-one, the wearable device can directly calculate the distances between the plurality of posterior probability vectors and corresponding template vectors and obtain the mean value. For example, in a case that the time length of the current voice source input speech is consistent with the time length of the user input speech during the wake-up word registration, the plurality of posterior probability vectors can correspond to the plurality of template vectors one-to-one. If the time length of the current voice source input speech is inconsistent with the time length of the user input speech during the wake-up word registration, the plurality of posterior probability vectors can not correspond to the plurality of template vectors one-to-one. In this case, the wearable device can establish a mapping relationship between the plurality of posterior probability vectors and the plurality of template vectors by using a dynamic time warping (DTW) method, so that the wearable device can calculate the distances between the plurality of posterior probability vectors and corresponding template vectors. That is, the wearable device can solve the template matching problem in a case of different lengths of data by using the DTW.
[0195] The first implementation manner of detecting the wake-up word based on the bone conduction signal by the wearable device is introduced above. In the first implementation manner, the wearable device first determines the fusion signal based on the bone conduction signal (including four manners), and then processes the fusion signal by using the acoustic model to obtain the posterior probability vector. Then, the wearable device decodes the obtained posterior probability vector based on the decoding path corresponding to the wake-up word to obtain the confidence degree corresponding to the command word currently input by the voice source. In a case that the confidence degree is greater than a confidence degree threshold, the wearable device determines that the command word including the wake-up word is detected. Alternatively, in a case that the confidence degree is greater than the confidence degree threshold and the obtained posterior probability vector matches the template vector, the wearable device determines that the command word including the wake-up word is detected. Next, the second implementation manner of detecting the wake-up word based on the bone conduction signal by the wearable device is introduced.
[0196] The second implementation manner
[0197] In the embodiments of the present application, before the wearable device detects the wake-up word based on the bone conduction signal, the air microphone is turned on, and the air conduction signal is collected through the air microphone. For example, in the case of detecting voice input, the air microphone is turned on, and the air conduction signal is collected through the air microphone. The wearable device determines a plurality of posterior probability vectors based on the bone conduction signal and the air conduction signal, and detects the wake-up word based on the plurality of posterior probability vectors. Among them, the plurality of posterior probability vectors one-to-one correspond to a plurality of audio frames included in the bone conduction signal and the air conduction signal, and a first posterior probability vector in the plurality of posterior probability vectors is used to indicate the probability that the phonemes in a first audio frame in the plurality of audio frames belong to a plurality of specified phonemes. It should be noted that the plurality of audio frames include the audio frames included in the bone conduction signal and the audio frames included in the air conduction signal. That is, each posterior probability vector in the plurality of posterior probability vectors corresponds to an audio frame included in the bone conduction signal or the air conduction signal, and a posterior probability vector indicates the probability that the phonemes in the corresponding audio frame belong to a plurality of specified phonemes.
[0198] It should be noted that the relevant introduction of the bone conduction signal and the air conduction signal can refer to the content in the foregoing first implementation mode, including the generation principle of the bone conduction signal and the air conduction signal, the preprocessing of the bone conduction signal and the air conduction signal, and the like, which will not be repeated here.
[0199] Next, first introduce the implementation mode that the wearable device determines a plurality of posterior probability vectors based on the bone conduction signal and the air conduction signal. It should be noted that there are many ways for the wearable device to determine a plurality of posterior probability vectors based on the bone conduction signal and the air conduction signal, and three of them are introduced next.
[0200] Way 1 of determining a plurality of posterior probability vectors based on a bone conduction signal and an air conduction signalThe wearable device inputs a start part of the bone conduction signal and the air conduction signal into the second acoustic model to obtain a first quantity of bone conduction posterior probability vectors and a second quantity of air conduction posterior probability vectors output by the second acoustic model. The start part of the bone conduction signal is determined according to a detection time delay of the voice detection, the first quantity of bone conduction posterior probability vectors correspond to audio frames included in the start part of the bone conduction signal one by one, and the second quantity of air conduction posterior probability vectors correspond to audio frames included in the air conduction signal one by one. The wearable device fuses a first bone conduction posterior probability vector and a first air conduction posterior probability vector to obtain a second posterior probability vector. The first bone conduction posterior probability vector corresponds to a last audio frame of the start part of the bone conduction signal, and a time length of the last audio frame is less than a frame time length. The first air conduction posterior probability vector corresponds to a first audio frame of the air conduction signal, and a time length of the first audio frame is less than the frame time length. The wearable device finally determines a plurality of posterior probability vectors including the second posterior probability vector, vectors other than the first bone conduction posterior probability vector in the first quantity of bone conduction posterior probability vectors, and vectors other than the first air conduction posterior probability vector in the second quantity of air conduction posterior probability vectors. The first quantity and the second quantity can be the same or different.
[0201] It should be noted that the foregoing first implementation mode can be referred to for the introduction of the start part of the bone conduction signal. In the embodiments of the present application, the last audio frame of the start part of the bone conduction signal can not be a complete audio frame, i.e., the time length of the last audio frame is less than the frame time length, for example, the start part of the bone conduction signal includes an audio frame with a half frame time length. Due to the loss of the head of the air conduction signal, the first audio frame of the air conduction signal can not be a complete audio frame, i.e., the time length of the first audio frame is less than the frame time length, for example, the first audio frame of the air conduction signal includes an audio frame with a half frame time length. In addition, the time length of the last audio frame of the start part of the bone conduction signal plus the time length of the first audio frame of the air conduction signal can equal the frame time length. In short, due to the voice detection (such as VAD), the start part of the bone conduction signal and the first frame of the air conduction signal can be incomplete, and the start part of the bone conduction signal and the first frame of the air conduction signal together represent the information of a complete audio frame. It should be noted that this complete audio frame is a potential frame of audio, not an actual frame. Alternatively, the wearable device adds the first bone conduction posterior probability vector and the first air conduction posterior probability vector to obtain the second posterior probability vector, and the second posterior probability vector obtained by the wearable device indicates the probability that the phoneme of the above-mentioned complete audio frame belongs to a plurality of specified phonemes.
[0202] That is, if the detection time delay of the voice detection is not an integer multiple of the frame duration, the duration of the last audio frame of the starting portion of the bone conduction signal is less than the frame duration, and the duration of the first audio frame of the air conduction signal is less than the frame duration, the wearable device needs to fuse (for example, add) the second bone conduction posterior probability vector and the second air conduction posterior probability vector, so as to obtain a plurality of posterior probability vectors. Alternatively, if the detection time delay of the voice detection is an integer multiple of the frame duration, the duration of the last audio frame of the starting portion of the bone conduction signal is equal to the frame duration, and the duration of the first audio frame of the air conduction signal is equal to the frame duration, the wearable device obtains the first number of bone conduction posterior probability vectors and the second number of air conduction posterior probability vectors as the plurality of posterior probability vectors, and then performs subsequent processing.
[0203] Figure 11 is another structure diagram of an acoustic model provided by an embodiment of the present application. Figure 11 The acoustic model shown is a second acoustic model in an embodiment of the present application. As can be seen, the second acoustic model in the embodiment of the present application includes two input layers (not shown), a shared network layer, and two output layers. The two input layers are used to input the starting portion of the bone conduction signal and the air conduction signal, respectively. The shared network layer is used to process the input data of the two input layers, respectively, to extract the features of the starting portion of the bone conduction signal and the air conduction signal, respectively. The two output layers are used to receive two output data of the shared network layer, respectively, and process the two output data, respectively, to output a first number of bone conduction posterior probability vectors corresponding to the starting portion of the bone conduction signal, and a second number of air conduction posterior probability vectors corresponding to the air conduction signal. That is, the wearable device processes the two parts of signals, the starting portion of the bone conduction signal and the air conduction signal, respectively, through the second acoustic model, to obtain two groups of posterior probability vectors corresponding to the two parts of signals. However, there is a shared network layer in the acoustic model, so that the two parts of signals share part of the network parameters.
[0204] In the embodiment of the present application, the wearable device fuses the first bone conduction posterior probability vector and the first air conduction posterior probability vector to obtain a second posterior probability vector, so that the plurality of bone conduction posterior probability vectors and the plurality of air conduction posterior probability vectors are fused, and then a plurality of posterior probability vectors are obtained, that is, the wearable device fuses the posterior probabilities of the two parts of signals, so that the plurality of posterior probability vectors obtained contain the command word information input by the sound source, which can also be regarded as a method of compensating for the loss of the air conduction signal based on the bone conduction signal, but not by directly fusing (such as splicing) the signals.
[0205] Way 2 of determining a plurality of posterior probability vectors based on a bone conduction signal and an air conduction signal The wearable device inputs the starting part of the bone conduction signal and the air conduction signal into the third acoustic model to obtain a plurality of posterior probability vectors output by the third acoustic model. It should be noted that the related introduction of the starting part of the bone conduction signal can also refer to the foregoing first implementation mode, and will not be described here.
[0206] In the embodiment of the present application, as shown in Figure 12 The third acoustic model includes two input layers (such as one input layer including DNN and CNN layers), a concatenation layer (concat layer), a network parameter layer (such as including RNN layer), and an output layer (such as including softmax layer). The two input layers are used to input the bone conduction signal and the air conduction signal respectively, the concatenation layer is used to splice the output data of the two input layers, the network parameter layer is used to process the output data of the concatenation layer, and the output layer is used to output a set of posterior probability vectors. That is, the wearable device inputs the starting part of the bone conduction signal and the air conduction signal into the third acoustic model, and implicitly fuses the starting part of the bone conduction signal and the air conduction signal through the concatenation layer in the third acoustic model, and then obtains a set of posterior probability vectors, so that the plurality of posterior probability vectors obtained contain the command word information input by the sound source, which can also be regarded as a method of compensating for the loss of the air conduction signal based on the bone conduction signal, but not by directly fusing the signals.
[0207] Way 3 of determining a plurality of posterior probability vectors based on a bone conduction signal and an air conduction signalThe wearable device inputs the bone conduction signal and the air conduction signal into the third acoustic model to obtain a plurality of posterior probability vectors output by the third acoustic model. That is, the wearable device directly inputs the bone conduction signal and the air conduction signal into the third acoustic model at the same time, and outputs a group of posterior probability vectors through the third acoustic model, so that the plurality of posterior probability vectors obtained contain the command word information input by the sound source, which can also be regarded as a method of compensating for the loss of the air conduction signal based on the bone conduction signal, but not by directly fusing the signals for compensation.
[0208] Next, the implementation of the wearable device detecting the wake-up word based on the plurality of posterior probability vectors is introduced.
[0209] In the embodiments of the present application, the wearable device determines the confidence of the phoneme sequence corresponding to the command word input by the sound source including the phoneme sequence corresponding to the wake-up word based on the plurality of posterior probability vectors and the phoneme sequence corresponding to the wake-up word. In the case where the confidence exceeds the confidence threshold, it is determined that the wake-up word is detected. The specific implementation is described in the foregoing first implementation, which will not be repeated here.
[0210] Optionally, in order to reduce the false wake-up rate, in the case where the confidence exceeds the confidence threshold and the distance condition is met between the plurality of posterior probability vectors and the plurality of template vectors, the wearable device determines that the command word including the wake-up word is detected. Optionally, in the case where the plurality of posterior probability vectors correspond to the plurality of template vectors one by one, the distance condition includes that the average distance between the plurality of posterior probability vectors and the corresponding template vectors is less than the distance threshold. The specific implementation is described in the foregoing first implementation, which will not be repeated here.
[0211] Step 403: When it is detected that the command word includes the wake-up word, the voice of the device to be woken up is woken up.
[0212] In the embodiments of the present application, when it is detected that the command word input by the sound source includes the wake-up word, the wearable device is woken up by voice. For example, the wearable device sends a wake-up instruction to the smart device (i.e. the device to be woken up) to wake up the smart device. Or, in the case where the wearable device itself is the smart device, the wearable device wakes up other components or modules except the bone conduction microphone, that is, the whole wearable device enters the working state.
[0213] As known from the above, the method for voice wake-up provided in the embodiments of the present application has a plurality of implementation manners, such as the first implementation manner and the second implementation manner described above, and each of the two implementation manners includes a plurality of specific implementation manners. Next, please refer to Figures 13 to 18 The above-described several specific implementations are explained again.
[0214] Figure 13This is a flowchart of another voice wake-up method provided in an embodiment of this application. Figure 13 This corresponds to Method 1 in the first implementation described above. Taking voice wake-up via multiple modules in a wearable device as an example, the wearable device acquires bone conduction signals through a bone conduction microphone, and performs VAD on the bone conduction signals through a VAD control module. When voice input is detected, the VAD control module outputs a high-level VAD control signal. When no voice input is detected, the VAD control module outputs a low-level VAD control signal. The VAD control module sends the VAD control signal to the air microphone control module, the front-end enhancement module, and the recognition engine, respectively. The VAD control signal is used to control the on / off states of the air microphone control module, the front-end enhancement module, and the recognition engine. When the VAD control signal is high, the air microphone control module controls the air microphone to turn on to acquire air conduction signals, the front-end enhancement module turns on to perform front-end enhancement on the air conduction signals, and the recognition engine turns on to detect wake-up words based on the bone conduction signals and air conduction signals. The fusion module performs preprocessing on the bone conduction signals, such as downsampling and / or gain adjustment, and uses the initial part of the preprocessed bone conduction signal to compensate for head drops in the front-end enhanced air conduction signal to obtain the fused signal. The fusion module sends the fused signal to the recognition engine, which uses a first acoustic model to identify the fused signal and obtain the detection result of the wake-up word. The recognition engine then sends the detection result to the processor (as shown in the diagram, the microcontroller unit (MCU)). The processor determines whether to wake up the smart device based on the detection result. If the detection result indicates that the input command word from the sound source includes a wake-up word, the processor wakes up the smart device via voice. If the detection result indicates that no wake-up word is detected, the processor does not wake up the smart device.
[0215] Figures 14 to 16 Here are flowcharts of three more voice wake-up methods provided in the embodiments of this application. Figure 14 , Figure 15 , Figure 16 respectively with Figure 13 The difference is that, in Figure 14 In the method shown, the fusion module generates an enhanced initiation signal based on the initial portion of the preprocessed bone conduction signal, and uses the enhanced initiation signal to perform head loss compensation on the enhanced air conduction signal at the front end to obtain the fused signal. Figure 15 In the method shown, the fusion module directly splices the preprocessed bone conduction signal and the enhanced air conduction signal at the front end to compensate for head loss in the air conduction signal, thereby obtaining the fused signal. Figure 16 In the method shown, the VAD control signal does not need to be sent to the air microphone control module, and therefore there is no need to collect the air conduction signal. In addition, the recognition engine directly determines the preprocessed bone conduction signal as the fusion signal.
[0216] Figure 17 is another method flowchart for voice wake-up provided by an embodiment of the present application. Figure 17 The difference between the method shown in Figure 13 and the method shown in Figure 17 is that, in the method shown in , the recognition engine inputs the preprocessed starting part of the bone conduction signal and the front-end enhanced air conduction signal into a second acoustic model respectively, obtains the bone conduction posterior probability vector and the air conduction posterior probability vector output respectively by two output layers of the second acoustic model, that is, obtains a posterior probability pair, and fuses the bone conduction posterior probability vector and the air conduction posterior probability vector to obtain a plurality of posterior probability vectors, and decodes the plurality of posterior probability vectors to obtain the detection result of the wake-up word.
[0217] Figure 18 is another method flowchart for voice wake-up provided by an embodiment of the present application. Figure 18 The difference between the method shown in Figure 17 and the method shown in Figure 18 is that, in the method shown in , the recognition engine inputs the preprocessed starting part of the bone conduction signal and the front-end enhanced air conduction signal into a third acoustic model respectively, or inputs the preprocessed bone conduction signal and the front-end enhanced air conduction signal into the third acoustic model respectively, and obtains a plurality of posterior probability vectors output respectively by one output layer of the third acoustic model.
[0218] From the above, in the embodiment of the present application, the bone conduction signal collected by the bone conduction microphone is used for voice detection, which can ensure low power consumption. While ensuring low power consumption, it is considered that the delay of voice detection may cause the collected air conduction signal to lose the head, so that the complete information of the command word input by the sound source is not contained, while the bone conduction signal collected by the bone conduction microphone contains the command word information input by the sound source, that is, the bone conduction signal does not lose the head. Therefore, the present scheme is based on the bone conduction signal to detect the wake-up word. In this way, the recognition accuracy of the wake-up word is high, and the precision of voice wake-up is high. In specific implementation, the bone conduction signal can be directly or implicitly compensated for head loss based on the bone conduction signal, or the wake-up word can be directly detected based on the bone conduction signal.
[0219] The above introduces the implementation process of voice wake-up based on the bone conduction signal of the wearable device. In the embodiment of the present application, the wake-up word can also be registered in the wearable device. Optionally, the confidence threshold in the above embodiment can also be determined at the same time of registering the wake-up word, and the plurality of template vectors in the above embodiment can also be determined. Next, the registration process of the wake-up word will be introduced.
[0220] In the embodiments of the present application, the wearable device first determines the phoneme sequence corresponding to the wake-up word. Then, the wearable device acquires the bone conduction registration signal, which contains complete information of the wake-up word. The wearable device determines the confidence threshold based on the bone conduction registration signal and the phoneme sequence corresponding to the wake-up word. Optionally, the wearable device can also determine a plurality of template vectors based on the bone conduction signal.
[0221] Optionally, the wearable device acquires the input wake-up word and determines the phoneme sequence corresponding to the wake-up word according to the pronunciation dictionary. Taking the case where the user inputs the wake-up word text to the wearable device as an example, the wearable device acquires the wake-up word text input by the user and determines the phoneme sequence corresponding to the wake-up word text according to the pronunciation dictionary. Optionally, in the process of registering the wake-up word, the wearable device can also detect whether the input wake-up word text meets the text registration condition after the user inputs the wake-up word text. If the wake-up word text meets the text registration condition, the wearable device determines the phoneme sequence corresponding to the wake-up word text according to the pronunciation dictionary.
[0222] For example, the text registration condition includes the text input frequency requirement and the character requirement. Taking the case where the text input frequency requirement requires the user to input the wake-up word text once or more times as an example, the wearable device performs text verification and analysis on the input wake-up word text each time the user inputs the wake-up word text to check whether the current input wake-up word text meets the character requirement. If the wake-up word text input by the user does not meet the character requirement, the wearable device prompts the user of the reason for not meeting the requirement and requires the user to re-input. If the wake-up word text input by the user one or more times all meet the character requirement and are the same, the wearable device determines the phoneme sequence corresponding to the wake-up word text according to the pronunciation dictionary.
[0223] Optionally, the wearable device detects whether the current input wake-up word text meets the character requirement through text verification. For example, the character requirement includes one or more of the following requirements: requiring Chinese (not Chinese is not meeting the character requirement), 4 to 6 characters (less than 4 characters or more than 6 characters is not meeting the character requirement), no mood auxiliary word (existence is not meeting the character requirement), no more than 3 repeated characters with the same pronunciation (existence is not meeting the character requirement), different from the existing command words (existence is not meeting the character requirement), the proportion of phonemes overlapping with the existing command words is not more than 70% (more than 70% is not meeting the character requirement, which is used to prevent false intrusion), the corresponding phonemes belong to the phonemes in the pronunciation dictionary (not belonging is not meeting the character requirement, which is an abnormal situation).
[0224] The above is the process of text registration, and the text registration can determine the phoneme sequence corresponding to the wake-up word. After the wearable device determines the phoneme sequence corresponding to the wake-up word, the phoneme sequence can be used as a decoding path of the wake-up word in subsequent processes, and the decoding path is used for detecting the wake-up word in the process of voice wake-up.
[0225] In addition to text registration, voice registration is also needed. In the embodiments of the present application, after the text registration is completed, the wearable device also needs to obtain a bone conduction registration signal containing complete information of the wake-up word. Optionally, the wearable device obtains an air conduction registration signal at the same time as obtaining the bone conduction registration signal. Optionally, in the process of voice registration, taking the wearable device obtaining the bone conduction registration signal and the air conduction registration signal input by the user as an example, after obtaining the input bone conduction registration signal and the air conduction registration signal, the wearable device checks whether the bone conduction registration signal and the air conduction registration signal meet the voice registration condition, and in the case of meeting the voice registration condition, the wearable device performs subsequent processing to determine the confidence threshold.
[0226] Exemplarily, the voice registration condition includes a voice input frequency requirement, a signal-to-noise ratio requirement, and a path score requirement. Taking the voice input frequency requirement as an example, the user needs to input the wake-up word voice (including the bone conduction registration signal and the air conduction registration signal) three times, the wearable device detects the wake-up word voice input by the user once, and then performs pronunciation verification and analysis on the input wake-up word voice to check whether the current input wake-up word voice meets the signal-to-noise ratio requirement and the path score requirement. If the wake-up word text input by the user does not meet the character requirement, the wearable device prompts the user of the reason for not meeting the requirement through text or sound and requires the user to re-input. If the wake-up word voice input by the user three times all meets the signal-to-noise ratio requirement and the path score requirement, the wearable device determines that the wake-up word voice input by the user meets the voice registration condition, and the wearable device performs subsequent processing.
[0227] Optionally, the wearable device can first detect whether the input wake-up word voice meets the signal-to-noise ratio requirement, and then detect whether the input wake-up word voice meets the path score requirement after determining that the input wake-up word voice meets the signal-to-noise ratio requirement. For example, the signal-to-noise ratio requirement includes a requirement that the signal-to-noise ratio is not lower than a signal-to-noise ratio threshold (lower than which does not meet the signal-to-noise ratio requirement), for example, a requirement that the signal-to-noise ratio of the bone conduction registration signal is not lower than a first signal-to-noise ratio threshold, and / or a requirement that the signal-to-noise ratio of the air conduction registration signal is not lower than a second signal-to-noise ratio threshold. Optionally, the first signal-to-noise ratio threshold is greater than the second signal-to-noise ratio threshold. If the user input wake-up word voice does not meet the signal-to-noise ratio requirement, the wearable device prompts the user that the current environment is too noisy and is not suitable for registration, and the user needs to find a quiet environment to re-input the wake-up word voice. The path score requirement includes a requirement that the path score obtained based on each input wake-up word voice is not less than a calibration threshold, a requirement that the average of three path scores obtained based on three input wake-up word voices is not less than the calibration threshold, and a requirement that the difference between any two path scores obtained based on two input wake-up word voices is not more than 100 points (or other values). The implementation process of obtaining the path score based on the wake-up word voice will be described below, which is similar to the process of obtaining the confidence level based on the bone conduction signal in the foregoing voice wake-up process.
[0228] Next, the implementation of the wearable device determining the confidence level threshold based on the bone conduction registration signal and the phoneme sequence corresponding to the wake-up word is introduced. Similar to the process of obtaining the confidence level based on the bone conduction signal in the foregoing voice wake-up process, the wearable device can determine the confidence level threshold in various implementation manners, two of which are introduced below.
[0229] First implementation manner
[0230] The wearable device determines the fusion registration signal based on the bone conduction registration signal, and determines the confidence level threshold and the plurality of template vectors based on the fusion registration signal and the phoneme sequence corresponding to the wake-up word.
[0231] First, the implementation of the wearable device determining the fusion registration signal based on the bone conduction registration signal is introduced. It should be noted that there are various ways for the wearable device to determine the fusion registration signal based on the bone conduction registration signal, four of which are introduced below.
[0232] Way 1 of determining a fusion registration signal based on a bone conduction registration signal Before determining the fusion registration signal based on the bone conduction registration signal, the air conduction registration signal is obtained. The wearable device fuses the starting part of the bone conduction registration signal and the air conduction registration signal to obtain the fusion registration signal. The starting part of the bone conduction registration signal is determined according to the detection time delay of the voice detection. Optionally, signal fusion is performed by signal splicing in the embodiments of the present application.
[0233] It should be noted that the wearable device fuses the starting part of the bone conduction registration signal and the air conduction registration signal, and the principle of the above-mentioned manner 1 of determining the fused signal based on the bone conduction signal is similar, and details are not introduced here. In addition, the wearable device can also detect the speech segment of the bone conduction registration signal and the air conduction registration signal to intercept the speech segment, and splice the signal based on the intercepted speech segment, thereby reducing the data processing amount. The bone conduction registration signal and the air conduction registration signal can also be preprocessed, for example, the bone conduction registration signal is down-sampled and / or gain-adjusted, and the air conduction signal is speech enhanced. The specific implementation manner is similar to the principle of the related content in the above-mentioned embodiments, please refer to the above-mentioned embodiments, and details are not introduced here.
[0234] Way 2 of determining a fusion registration signal based on a bone conduction registration signal Before determining the fused registration signal based on the bone conduction registration signal, the air conduction registration signal is obtained. The wearable device generates an enhanced starting registration signal based on the starting part of the bone conduction registration signal, and fuses the enhanced starting registration signal and the air conduction registration signal to obtain the fused registration signal. Wherein, the starting part of the bone conduction registration signal is determined according to the detection time delay of the speech detection.
[0235] It should be noted that the difference between the above-mentioned manner 1 of determining the fused registration signal based on the bone conduction registration signal and the above-mentioned manner 2 is that in the manner 2, the wearable device generates an enhanced starting registration signal based on the starting part of the bone conduction registration signal, and fuses the enhanced starting registration signal and the air conduction registration signal, rather than fusing the starting part of the bone conduction registration signal and the air conduction signal. In addition, in the manner 2, the wearable device can also detect the speech segment of the bone conduction registration signal and the air conduction registration signal to intercept the speech segment, and fuse the signal based on the intercepted speech segment, thereby reducing the data processing amount. The wearable device can also preprocess the bone conduction registration signal and the air conduction registration signal, for example, down-sample and / or gain-adjust the bone conduction registration signal, and speech enhance the air conduction signal. The specific implementation manner is similar to the principle of the related content in the above-mentioned embodiments, please refer to the above-mentioned embodiments, and details are not introduced here.
[0236] In the embodiments of the present application, the wearable device can input the starting part of the bone conduction registration signal into a generation network model to obtain an enhanced starting registration signal output by the generation network model. Wherein, the generation network model can be the same as the generation network model introduced above, or can be another generation network model, and the present application is not limited to this. The network structure, training method, training device, etc. of the generation network model are not limited in the present application.
[0237] Way 3 of determining a fusion registration signal based on a bone conduction registration signal: The wearable device obtains the air conduction registration signal before determining the fusion registration signal based on the bone conduction registration signal. The wearable device directly fuses the bone conduction registration signal and the air conduction registration signal to obtain the fusion registration signal.
[0238] It should be noted that the difference between the above-mentioned mode 1 of determining the fusion registration signal based on the bone conduction registration signal is that, in this mode 3, the wearable device directly fuses the bone conduction registration signal and the air conduction registration signal to obtain the fusion registration signal. In addition, in this mode 3, the wearable device can also perform speech segment detection on the bone conduction registration signal and the air conduction registration signal to intercept the speech segment, and perform signal fusion based on the intercepted speech segment, thereby reducing the data processing amount. The wearable device can also preprocess the bone conduction registration signal and the air conduction registration signal, such as downsampling and / or gain adjustment on the bone conduction registration signal, speech enhancement on the air conduction signal, etc. The specific implementation mode is similar to the principle of the related content in the foregoing embodiments, please refer to the foregoing embodiments, which will not be described in detail here.
[0239] Way 4 of determining a fusion registration signal based on a bone conduction registration signal : The wearable device determines the bone conduction registration signal as the fusion registration signal.
[0240] It should be noted that the difference between the above-mentioned mode 1 of determining the fusion registration signal based on the bone conduction registration signal is that, in this mode 4, the wearable device directly uses the bone conduction registration signal as the fusion registration signal. In addition, in this mode 4, the wearable device can also perform speech segment detection on the bone conduction registration signal to intercept the speech segment, and perform subsequent processing based on the intercepted speech segment, thereby reducing the data processing amount. The wearable device can also preprocess the bone conduction registration signal, such as downsampling and / or gain adjustment on the bone conduction registration signal, etc. The specific implementation mode is similar to the principle of the related content in the foregoing embodiments, please refer to the foregoing embodiments, which will not be described in detail here.
[0241] Next, the implementation mode of the wearable device determining the confidence threshold and the plurality of template vectors based on the fusion registration signal and the phoneme sequence corresponding to the wake-up word is introduced.
[0242] Optionally, the wearable device inputs the plurality of enrollment audio frames included in the fusion enrollment signal into the first acoustic model to obtain a plurality of enrollment posterior probability vectors output by the first acoustic model. The plurality of enrollment posterior probability vectors correspond to the plurality of enrollment audio frames one-to-one, and a first enrollment posterior probability vector in the plurality of enrollment posterior probability vectors indicates a probability that a phoneme of a first enrollment audio frame in the plurality of enrollment audio frames belongs to the plurality of specified phonemes. That is, each enrollment posterior probability vector in the plurality of enrollment posterior probability vectors corresponds to an enrollment audio frame included in the fusion enrollment signal, and the enrollment posterior probability vector indicates a probability that a phoneme of the corresponding enrollment audio frame belongs to the plurality of specified phonemes. The wearable device determines the plurality of enrollment posterior probability vectors as the plurality of template vectors. The wearable device determines the confidence threshold based on the plurality of enrollment posterior probability vectors and the phoneme sequence corresponding to the wake-up word. That is, the wearable device processes the fusion enrollment signal through the first acoustic model to obtain information of the phonemes included in the fusion signal, that is, the enrollment posterior probability vectors, takes the enrollment posterior probability vectors as the template vectors, and stores the template vectors. The wearable device also decodes the enrollment posterior probability vectors based on the phoneme sequence corresponding to the wake-up word (that is, the decoding path) to determine a path score, takes the path score as the confidence threshold, and stores the confidence threshold. Details of the first acoustic model are described above and will not be described here.
[0243] The above describes a first implementation of the wearable device determining the confidence threshold based on the bone conduction enrollment signal and the phoneme sequence corresponding to the wake-up word. In the first implementation, the wearable device first determines the fusion enrollment signal based on the bone conduction enrollment signal (including four ways), and then processes the fusion enrollment signal through the acoustic model to obtain the enrollment posterior probability vectors. Then, the wearable device decodes the obtained enrollment posterior probability vectors based on the decoding path corresponding to the wake-up word to obtain the confidence threshold. Optionally, the wearable device stores the obtained enrollment posterior probability vectors as the template vectors. Next, a second implementation of the wearable device determining the confidence threshold based on the bone conduction enrollment signal and the phoneme sequence corresponding to the wake-up word is described.
[0244] Second implementation
[0245] In the embodiment of the present application, before determining the confidence threshold based on the bone conduction registration signal and the phoneme sequence corresponding to the wake-up word, the wearable device acquires an air conduction registration signal. The wearable device determines a plurality of registration posterior probability vectors based on the bone conduction registration signal and the air conduction registration signal. The plurality of registration posterior probability vectors correspond one-to-one to a plurality of registration audio frames included in the bone conduction registration signal and the air conduction registration signal. A first registration posterior probability vector in the plurality of registration posterior probability vectors indicates a probability that a phoneme of a first registration audio frame in the plurality of registration audio frames belongs to the plurality of specified phonemes. It should be noted that the plurality of registration audio frames include registration audio frames included in the bone conduction registration signal and registration audio frames included in the air conduction registration signal. That is, each registration posterior probability vector in the plurality of registration posterior probability vectors corresponds to one registration audio frame included in the bone conduction registration signal or the air conduction registration signal, and one registration posterior probability vector indicates a probability that a phoneme of the corresponding one registration audio frame belongs to the plurality of specified phonemes. The wearable device determines the confidence threshold based on the plurality of registration posterior probability vectors and the phoneme sequence corresponding to the wake-up word. Optionally, the wearable device determines the plurality of registration posterior probability vectors as a plurality of template vectors.
[0246] It should be noted that the relevant introduction of the bone conduction registration signal and the air conduction registration signal can refer to the content in the foregoing first implementation mode, including the generation principle of the bone conduction registration signal and the air conduction registration signal, the preprocessing of the bone conduction registration signal and the air conduction registration signal, and the like, which will not be repeated here.
[0247] Next, first introduce the implementation mode that the wearable device determines a plurality of registration posterior probability vectors based on the bone conduction registration signal and the air conduction registration signal. It should be noted that there are many ways for the wearable device to determine a plurality of registration posterior probability vectors based on the bone conduction registration signal and the air conduction registration signal, and three of them are introduced next.
[0248] Way 1 of determining a plurality of registration posterior probability vectors based on a bone conduction registration signal and an air conduction registration signal: The wearable device inputs the initial part of the bone conduction registration signal and the air conduction registration signal into the second acoustic model to obtain third quantity of bone conduction registration posterior probability vectors and fourth quantity of air conduction registration posterior probability vectors output by the second acoustic model. The wearable device fuses the first bone conduction registration posterior probability vector and the first air conduction registration posterior probability vector to obtain a second registration posterior probability vector. The initial part of the bone conduction registration signal is determined according to a detection time delay of the speech detection. The third quantity of bone conduction registration posterior probability vectors correspond to registration audio frames included in the initial part of the bone conduction registration signal one by one. The fourth quantity of air conduction registration posterior probability vectors correspond to registration audio frames included in the air conduction registration signal one by one. That is, one bone conduction registration posterior probability vector corresponds to one registration audio frame included in the initial part of the bone conduction registration signal, and one air conduction registration posterior probability vector corresponds to one registration audio frame included in the air conduction registration signal. The first bone conduction registration posterior probability vector corresponds to a last registration audio frame of the initial part of the bone conduction registration signal, and a time length of the last registration audio frame is less than a frame time length. The first air conduction registration posterior probability vector corresponds to a first registration audio frame of the air conduction registration signal, and a time length of the first registration audio frame is less than the frame time length. The plurality of registration posterior probability vectors finally determined by the wearable device include the second registration posterior probability vector, vectors other than the first bone conduction registration posterior probability vector in the third quantity of bone conduction registration posterior probability vectors, and vectors other than the first air conduction registration posterior probability vector in the fourth quantity of air conduction registration posterior probability vectors. The third quantity and the fourth quantity can be the same or different. The third quantity and the first quantity described above can be the same or different. The fourth quantity and the second quantity described above can be the same or different.
[0249] Optionally, the wearable device adds the first bone conduction registration posterior probability vector and the first air conduction registration posterior probability vector to obtain the second registration posterior probability vector.
[0250] It should be noted that the related introduction about the initial part of the bone conduction registration signal can also refer to the content in the foregoing first implementation manner, which will not be described here again. In addition, the related introduction about the second acoustic model can refer to the related content of the foregoing embodiments, which will not be described here again. The principle of obtaining the third quantity of bone conduction registration posterior probability vectors and the fourth quantity of air conduction registration posterior probability vectors by the wearable device through the second acoustic model is consistent with the principle of obtaining the first quantity of bone conduction posterior probability vectors and the second quantity of air conduction posterior probability vectors by the wearable device through the second acoustic model in the foregoing embodiments, which will not be described in detail here.
[0251] Way 2 of determining a plurality of registration posterior probability vectors based on a bone conduction registration signal and an air conduction registration signal : The wearable device inputs the initial part of the bone conduction registration signal and the air conduction registration signal into the third acoustic model to obtain a plurality of registration posterior probability vectors output by the third acoustic model. The initial part of the bone conduction registration signal is determined according to a detection time delay of the speech detection.
[0252] It should be noted that the description of the starting part of the bone conduction registration signal can also refer to the foregoing first implementation mode, and will not be described here. In addition, the description of the third acoustic model can refer to the foregoing embodiment, and will not be described here. The principle of obtaining the plurality of registration posterior probability vectors by the wearable device through the third acoustic model is consistent with the principle of obtaining the plurality of posterior probability vectors through the third acoustic model in the foregoing embodiment, and will not be described here in detail.
[0253] Way 3 of determining a plurality of registration posterior probability vectors based on a bone conduction registration signal and an air conduction registration signal The wearable device inputs the bone conduction registration signal and the air conduction registration signal into the third acoustic model to obtain a plurality of registration posterior probability vectors output by the third acoustic model. That is, the wearable device directly inputs the bone conduction registration signal and the air conduction registration signal into the third acoustic model at the same time, and outputs a set of registration posterior probability vectors through the third acoustic model, so that the plurality of registration posterior probability vectors obtained contain complete information of the wake-up word input by the sound source.
[0254] In the embodiment of the present application, after the wearable device determines the plurality of registration posterior probability vectors, the confidence threshold is determined based on the plurality of registration posterior probability vectors and the phoneme sequence corresponding to the wake-up word. The principle is similar to the principle of determining the confidence of the wearable device based on the plurality of posterior probability vectors and the phoneme sequence of the wake-up word as described above, and the specific implementation mode can refer to the foregoing related description, which will not be described here in detail.
[0255] Figures 19 to 24 is the flow chart of the six wake-up word registration methods provided by the embodiment of the present application. Next, the registration process of the wake-up word in the embodiment of the present application will be explained again. Figures 19 to 24 The registration process of the wake-up word in the embodiment of the present application will be explained again.
[0256] Figure 19Corresponding to the above-mentioned manner 1 in the first implementation manner of the wake-up word registration, the registration process of the wake-up word includes text registration and voice registration. Taking the wake-up word registration of the wearable device through multiple modules as an example, the wearable device first performs text registration. The text registration module of the wearable device acquires the text of the wake-up word input by the user, performs text verification and text analysis on the input text of the wake-up word, and determines the phoneme sequence corresponding to the text of the wake-up word meeting the text registration requirement according to the pronunciation dictionary, and determines the phoneme sequence as a decoding path. The text registration module sends the decoding path to the recognition engine. The recognition engine stores the decoding path. The wearable device then performs voice registration. The voice registration module of the wearable device acquires the voice registration signal, including the bone conduction registration signal and the air conduction registration signal. Optionally, the wearable device acquires the bone conduction registration signal and the air conduction registration signal through VAD, and can also preprocess the acquired bone conduction registration signal and air conduction registration signal. Then, the voice registration module performs pronunciation verification on the bone conduction registration signal and the air conduction registration signal, and the fusion module fuses the bone conduction registration signal and the air conduction registration signal meeting the voice registration requirement after verification to obtain a fusion registration signal. In order to distinguish Figures 19 to 22 , the fusion registration signal in Figure 19 is referred to as fusion registration signal 1. The voice registration module processes the fusion registration signal 1 through the first acoustic model to obtain a plurality of registration posterior probability vectors, and decodes the plurality of registration posterior probability vectors to determine a path score, which is sent to the recognition engine as a wake-up threshold (i.e. confidence threshold). The recognition engine stores the wake-up threshold, which is used for primary false wake-up suppression in subsequent voice wake-up of the user. Optionally, the voice registration module sends the plurality of registration posterior probability vectors obtained as a plurality of template vectors to the recognition engine, and the recognition engine stores the plurality of template vectors, which are used for secondary false wake-up suppression in subsequent voice wake-up.
[0257] Figures 20 to 22 Corresponding to the above-mentioned manner 2, manner 3 and manner 4 in the first implementation manner of the wake-up word registration respectively. The difference between Figure 19 is that, in the method shown in Figure 20 , the voice registration module of the wearable device generates an enhanced initial registration signal based on the initial part of the bone conduction registration signal, fuses the enhanced initial registration signal with the air conduction registration signal to obtain a fusion registration signal. The fusion registration signal in Figure 20 is referred to as fusion registration signal 2. In the method shown in Figure 21 , the voice registration module directly fuses the bone conduction registration signal and the air conduction registration signal to obtain a fusion registration signal. The fusion registration signal in Figure 21 is referred to as fusion registration signal 3. In the method shown in Figure 22In the method shown, the voice registration module can directly determine the bone conduction registration signal as the fusion registration signal without acquiring the air conduction registration signal. Here, we will... Figure 22 The fusion registration signal in the above is called fusion registration signal 4.
[0258] Figure 23 This corresponds to Method 1 in the second implementation of the wake word registration described above. (And...) Figure 19 The difference is that, in Figure 23 In the method shown, the voice registration module of the wearable device inputs the initial portion of the bone conduction registration signal and the air conduction registration signal into the second acoustic model, respectively, to obtain a third number of bone conduction registration posterior probability vectors and a fourth number of air conduction registration posterior probability vectors output by the second acoustic model. The voice registration module fuses the third number of bone conduction registration posterior probability vectors and the fourth number of air conduction registration posterior probability vectors to obtain multiple registration posterior probability vectors.
[0259] Figure 24 This corresponds to methods 2 and 3 in the second implementation of wake word registration mentioned above. (And...) Figure 23 The difference is that, in Figure 24 In the method shown, the voice registration module of the wearable device inputs the initial part of the bone conduction registration signal and the air conduction registration signal into the third acoustic model, or inputs the bone conduction registration signal and the air conduction registration signal into the third acoustic model, so as to obtain multiple registration posterior probability vectors output by the third acoustic model.
[0260] As described above, the processing flow of bone conduction and air conduction registration signals during wake-up word registration is similar to that during voice wake-up. The difference lies in the approach: in wake-up word registration, the goal is to obtain the wake-up threshold and template vector, while in voice wake-up, the goal is to detect the wake-up word. The template vector enhances the accuracy and robustness of this scheme. This scheme directly or implicitly compensates for head loss in air conduction signals using bone conduction signals, or directly detects wake-up words based on bone conduction signals. Since bone conduction signals contain the command word information from the sound source input (i.e., no head loss occurs), the accuracy of wake-up word recognition and voice wake-up is high.
[0261] The above embodiments describe the voice wake-up process and the wake-up word registration process. As mentioned above, the acoustic models in this application embodiment need to be pre-trained, such as the first acoustic model, the second acoustic model, and the third acoustic model. The training process of the acoustic model will be described next using the example of training the acoustic model with a computer device.
[0262] In the embodiments of the present application, the computer device first acquires a second training data set, the second training data set comprising a plurality of second sample signal pairs, one second sample signal pair comprising one bone conduction sample signal and one air conduction sample signal, and one second sample signal pair corresponding to one command word. Optionally, the second training data set comprises directly collected speech data, public speech data and / or speech data purchased from a third party. Optionally, before training, the computer device can pre-process the acquired second training data set pairs to obtain a pre-processed second training data set, the pre-processed second training data set being able to simulate the distribution of real speech data so as to be closer to real scene speech and increase the diversity of training samples. Illustratively, the second training data set is backed up, i.e. an additional data is added, and the backed up data is pre-processed. Optionally, the backed up data is divided into multiple parts, each part of data is pre-processed in one way, and the pre-processing of each part of data can be different, so that the total training data is doubled, and the comprehensiveness of the data is ensured, a balance between performance and training cost is achieved, and the accuracy and robustness of speech recognition are improved to a certain extent. Wherein, the method of pre-processing each part of data can comprise one or more of adding noise, volume enhancement, adding reverberation, time shift, changing tone, time stretching, etc.
[0263] Taking the trained first acoustic model as an example, the computer device determines multiple fusion sample signals based on the second training dataset in four ways. It should be noted that these four methods correspond one-to-one with the four methods used by the wearable device to determine the fusion signal based on the bone conduction signal during the recognition process (i.e., the voice wake-up process) in the above embodiments. That is, if the wearable device fuses the initial part of the bone conduction signal and the air conduction signal during the recognition process to obtain a fusion signal, then during the training process, the computer device fuses the initial part of the bone conduction sample signal and the air conduction sample signal included in each of the multiple second sample signal pairs to obtain multiple fusion sample signals. If, during the recognition process, the wearable device generates an enhanced starting signal based on the initial part of the bone conduction signal and fuses the enhanced starting signal with the air conduction signal to obtain a fusion signal, then during the training process, the computer device generates an enhanced starting sample signal based on the initial part of the bone conduction sample signal included in each of the multiple second sample signal pairs and fuses each enhanced starting sample signal with the corresponding air conduction sample signal to obtain multiple fusion sample signals. If, during the recognition process, the computer device directly fuses the bone conduction signal and the air conduction signal to obtain a fused signal, then during the training process, the computer device directly fuses the bone conduction sample signal and the air conduction sample signal included in each of the multiple second sample signal pairs to obtain multiple fused sample signals. If, during the recognition process, the wearable device determines the bone conduction signal as the fused signal, then during the training process, the computer device determines the bone conduction sample signal included in the multiple second sample signal pairs as multiple fused sample signals. The starting portion of the bone conduction sample signal is determined based on the detection delay of speech detection or set empirically. Subsequently, the computer device trains a first initial acoustic model using these multiple fused sample signals to obtain the first acoustic model in this embodiment. The network structure of the first initial acoustic model is the same as the network structure of the first acoustic model.
[0264] Optionally, before determining multiple fused sample signals based on the second training dataset, the computer device preprocesses the bone conduction sample signals and air conduction sample signals included in the second training dataset, such as performing front-end enhancement on the air conduction sample signals and downsampling and gain adjustment on the bone conduction sample signals. Optionally, the computer device inputs the initial portion of the bone conduction sample signal included in each of the multiple second sample signal pairs into a generator network model to obtain the enhanced initial sample signal output by the generator network model. This generator network model can be the same model as the generator network model in the aforementioned embodiments, or it can be a different model. This application embodiment does not limit this.
[0265] For example, Figures 25 to 28are four schematic diagrams for training the first acoustic model based on the four manners respectively provided in the embodiments of the present application. Referring to Figures 25 to 28 The second training data set obtained by the computer device includes bone conduction data (bone conduction sample signals) and air conduction data (air conduction sample signals). The computer device performs downsampling and / or gain adjustment on the bone conduction data through the fusion module and performs front-end enhancement on the air conduction data through the front-end enhancement module. Figures 25 to 27 Corresponding to the first three manners of the four manners, the fusion module performs head loss compensation on the air conduction signal through the bone conduction data in a corresponding manner to obtain training input data. Figure 28 Corresponding to the fourth manner of the four manners, the fusion module directly takes the bone conduction data as the training input data without the air conduction data. Then, the computer device trains the network model (i.e., the first initial acoustic model) through the training input data and adjusts the network model through the loss function, the gradient descent algorithm and error back propagation to obtain the trained first acoustic model.
[0266] Taking the training of the second acoustic model as an example, corresponding to the manner 1 in which the wearable device determines a plurality of posterior probability vectors based on the bone conduction signal and the air conduction signal in the process of voice wake-up, in the training process, the computer device takes the starting part of the bone conduction sample signal and the air conduction sample signal included in each second sample signal pair in the plurality of second sample signal pairs as the input of the second initial acoustic model to train the second initial acoustic model to obtain the second acoustic model. The network structure of the second initial acoustic model is the same as that of the second acoustic model. That is, the second initial acoustic model also includes two input layers, one shared network layer and two output layers.
[0267] Figure 29 is a schematic diagram for training the second acoustic model provided in the embodiments of the present application. Referring to Figure 29 The second training data set obtained by the computer device includes bone conduction data (bone conduction sample signals) and air conduction data (air conduction sample signals). The computer device performs downsampling and / or gain adjustment on the bone conduction data and performs front-end enhancement on the air conduction data. The computer device takes the bone conduction data as the training input data 1 and takes the air conduction data as the training input data 2. The computer device trains the network model (i.e., the second initial acoustic model) through the training input data 1 and the training input data 2 and adjusts the network model through the loss function, the gradient descent algorithm and error back propagation to obtain the trained second acoustic model. The training input data 1 and the training input data 2 can correspond to the same loss function or different loss functions, which is not limited in the embodiments of the present application.
[0268] Corresponding to the manner 2 of determining the plurality of posterior probability vectors based on the bone conduction signal and the air conduction signal in the process of voice wake-up, in the training process, the computer device takes the starting part of the bone conduction sample signal and the air conduction sample signal included in each second sample signal pair in the plurality of second sample signal pairs as the input of the third initial acoustic model to train the third initial acoustic model, and obtains the third acoustic model. Alternatively, corresponding to the manner 3 of determining the plurality of posterior probability vectors based on the bone conduction signal and the air conduction signal in the process of voice wake-up, in the training process, the computer device takes the bone conduction sample signal and the air conduction sample signal included in each second sample signal pair in the plurality of second sample signal pairs as the input of the third initial acoustic model to train the third initial acoustic model, and obtains the third acoustic model. The network structure of the third initial acoustic model is the same as that of the third acoustic model. That is, the third initial acoustic model also includes two input layers, one splicing layer, one network parameter layer, and one output layer.
[0269] Exemplarily, Figure 30 is a schematic diagram of training a third acoustic model provided by an embodiment of the present application. Referring to Figure 30 , the second training data set obtained by the computer device includes bone conduction data (bone conduction sample signal) and air conduction data (air conduction sample signal). The computer device performs downsampling and / or gain adjustment on the bone conduction data, and performs front-end enhancement on the air conduction data. The computer device takes the bone conduction data or the starting part in the bone conduction data as training input data 1, and takes the air conduction data as training input data 2. The computer device trains a network model (i.e., a third initial acoustic model) through the training input data 1 and the training input data 2, and adjusts the network model through a loss function, and a gradient descent algorithm and error back propagation, thereby obtaining a trained third acoustic model.
[0270] In summary, in the training process, the bone conduction sample signal is also directly or implicitly used to compensate for the loss of the air conduction registration signal, thereby constructing training input data to train the initial acoustic model and obtaining the trained acoustic model. In the process of voice wake-up, the bone conduction signal is directly or implicitly compensated for the loss of the air conduction signal in the same way. Since the bone conduction signal contains the command word information of the sound source input, i.e., the bone conduction signal has no loss of head, the recognition accuracy of the detection of the wake-up word based on the bone conduction signal is high, the accuracy of voice wake-up is high, and the robustness is also improved.
[0271] Figure 31 is a structural schematic diagram of a voice wake-up device 3100 provided by an embodiment of the present application. The voice wake-up device 3100 can be realized by software, hardware, or a combination of both as part of or all of an electronic device, which can beFigure 2 The wearable device shown in FIG. 1. Referring to Figure 31 The device 3100 includes a voice detection module 3101, a wake-up word detection module 3102, and a voice wake-up module 3103.
[0272] The voice detection module 3101 is configured to perform voice detection based on a bone conduction signal collected by a bone conduction microphone, the bone conduction signal containing command word information input by a sound source.
[0273] The wake-up word detection module 3102 is configured to, when it is detected that there is voice input, perform wake-up word detection based on the bone conduction signal.
[0274] The voice wake-up module 3103 is configured to, when it is detected that the command word contains a wake-up word, perform voice wake-up on a device to be woken up.
[0275] Optionally, the wake-up word detection module 3102 includes:
[0276] A first determination sub-module configured to determine a fusion signal based on the bone conduction signal.
[0277] A wake-up word detection sub-module configured to perform wake-up word detection on the fusion signal.
[0278] Optionally, the device 3100 further includes:
[0279] A processing module configured to turn on an air microphone and collect an air conduction signal through the air microphone.
[0280] The first determination sub-module is configured to:
[0281] fuse a starting portion of the bone conduction signal and the air conduction signal to obtain the fusion signal, the starting portion of the bone conduction signal being determined according to a detection time delay of the voice detection; or
[0282] generate an enhanced starting signal based on the starting portion of the bone conduction signal, and fuse the enhanced starting signal and the air conduction signal to obtain the fusion signal, the starting portion of the bone conduction signal being determined according to the detection time delay of the voice detection; or
[0283] directly fuse the bone conduction signal and the air conduction signal to obtain the fusion signal.
[0284] Optionally, the wake-up word detection sub-module is configured to:
[0285] input a plurality of audio frames included in the fusion signal into a first acoustic model to obtain a plurality of posterior probability vectors output by the first acoustic model, the plurality of posterior probability vectors corresponding one-to-one to the plurality of audio frames, a first posterior probability vector in the plurality of posterior probability vectors being used to indicate a probability that a phoneme of a first audio frame in the plurality of audio frames belongs to a plurality of specified phonemes.
[0286] detect the wake-up word based on the plurality of posterior probability vectors.
[0287] Optionally, the apparatus 3100 further includes:
[0288] a processing module, configured to start the air microphone, and collect the air signal through the air microphone;
[0289] The wake-up word detection module 3102 includes:
[0290] a second determination sub-module, configured to determine, based on the bone signal and the air signal, a plurality of posterior probability vectors, the plurality of posterior probability vectors corresponding to a plurality of audio frames included in the bone signal and the air signal, a first posterior probability vector in the plurality of posterior probability vectors being used to indicate a probability that a phoneme in a first audio frame in the plurality of audio frames belongs to a plurality of specified phonemes;
[0291] a wake-up word detection sub-module, configured to detect the wake-up word based on the plurality of posterior probability vectors.
[0292] Optionally, the second determination sub-module is configured to:
[0293] input a starting portion of the bone signal and the air signal into the second acoustic model to obtain a first number of bone posterior probability vectors and a second number of air posterior probability vectors output by the second acoustic model, the starting portion of the bone signal being determined according to a detection time delay of the speech detection, the first number of bone posterior probability vectors corresponding to audio frames included in the starting portion of the bone signal, and the second number of air posterior probability vectors corresponding to audio frames included in the air signal;
[0294] fuse a first bone posterior probability vector and a first air posterior probability vector to obtain a second posterior probability vector, the first bone posterior probability vector corresponding to a last audio frame of the starting portion of the bone signal, a time length of the last audio frame being less than a frame time length, the first air posterior probability vector corresponding to a first audio frame of the air signal, a time length of the first audio frame being less than the frame time length, and the plurality of posterior probability vectors including the second posterior probability vector, vectors other than the first bone posterior probability vector in the first number of bone posterior probability vectors, and vectors other than the first air posterior probability vector in the second number of air posterior probability vectors.
[0295] Optionally, the second determination sub-module is configured to:
[0296] input the starting portion of the bone signal and the air signal into a third acoustic model to obtain the plurality of posterior probability vectors output by the third acoustic model, the starting portion of the bone signal being determined according to the detection time delay of the speech detection; or
[0297] The bone conduction signal and the air conduction signal are input into the third acoustic model to obtain a plurality of posterior probability vectors output by the third acoustic model.
[0298] Optionally, the wake-up word detection submodule is configured to:
[0299] Based on the plurality of posterior probability vectors and the phoneme sequence corresponding to the wake-up word, the confidence that the phoneme sequence corresponding to the command word includes the phoneme sequence corresponding to the wake-up word is determined.
[0300] In a case where the confidence exceeds a confidence threshold, it is determined that the command word including the wake-up word is detected.
[0301] Optionally, the wake-up word detection submodule is configured to:
[0302] Based on the plurality of posterior probability vectors and the phoneme sequence corresponding to the wake-up word, the confidence that the phoneme sequence corresponding to the command word includes the phoneme sequence corresponding to the wake-up word is determined.
[0303] In a case where the confidence exceeds a confidence threshold and the plurality of posterior probability vectors and the plurality of template vectors satisfy a distance condition, it is determined that the command word including the wake-up word is detected, the plurality of template vectors indicating probabilities that phonemes of the speech signal containing complete information of the wake-up word belong to a plurality of specified phonemes.
[0304] Optionally, in a case where the plurality of posterior probability vectors and the plurality of template vectors are in one-to-one correspondence, the distance condition includes that a mean value of distances between the plurality of posterior probability vectors and corresponding template vectors is less than a distance threshold.
[0305] Optionally, the apparatus 3100 further includes:
[0306] The obtaining module is configured to obtain a bone conduction registration signal containing complete information of the wake-up word.
[0307] The determining module is configured to determine a confidence threshold and a plurality of template vectors based on the bone conduction registration signal and the phoneme sequence corresponding to the wake-up word.
[0308] Optionally, the determining module includes:
[0309] The third determining submodule is configured to determine a fusion registration signal based on the bone conduction registration signal.
[0310] The fourth determining submodule is configured to determine a confidence threshold and a plurality of template vectors based on the fusion registration signal and the phoneme sequence corresponding to the wake-up word.
[0311] Optionally, the fourth determining submodule is configured to:
[0312] input the plurality of registration audio frames comprised in the fusion registration signal into the first acoustic model to obtain a plurality of registration posterior probability vectors output by the first acoustic model, the plurality of registration posterior probability vectors corresponding to the plurality of registration audio frames in a one-to-one manner, and a first registration posterior probability vector in the plurality of registration posterior probability vectors indicating a probability that a phoneme of a first registration audio frame in the plurality of registration audio frames belongs to the plurality of specified phonemes;
[0313] determine the plurality of registration posterior probability vectors as a plurality of template vectors;
[0314] determine a confidence threshold based on the plurality of registration posterior probability vectors and a phoneme sequence corresponding to the wake-up word.
[0315] Optionally, the apparatus 3100 further includes:
[0316] The obtaining module is configured to obtain an air-conduction registration signal.
[0317] The determining module includes:
[0318] The fifth determining sub-module is configured to determine, based on the bone-conduction registration signal and the air-conduction registration signal, a plurality of registration posterior probability vectors corresponding to a plurality of registration audio frames comprised in the bone-conduction registration signal and the air-conduction registration signal, and a first registration posterior probability vector in the plurality of registration posterior probability vectors indicating a probability that a phoneme of a first registration audio frame in the plurality of registration audio frames belongs to the plurality of specified phonemes.
[0319] The sixth determining sub-module is configured to determine a confidence threshold based on the plurality of registration posterior probability vectors and a phoneme sequence corresponding to the wake-up word.
[0320] In the embodiments of the present application, the bone-conduction microphone is used to collect the bone-conduction signal for voice detection, which can ensure low power consumption. In addition, considering that the delay of voice detection may cause the collected air-conduction signal to lose the head, thereby not containing the complete information of the command word input by the sound source, while the bone-conduction signal collected by the bone-conduction microphone contains the command word information input by the sound source, that is, the bone-conduction signal does not lose the head. Therefore, the present scheme is based on the bone-conduction signal to detect the wake-up word. In this way, the recognition accuracy of the wake-up word is high, and the accuracy of voice wake-up is high.
[0321] It should be noted that: the voice wake-up apparatus provided in the above embodiments is only used as an example for illustrating the division of the above functional modules in the voice wake-up process. In actual applications, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the apparatus is divided into different functional modules to complete all or part of the functions described above. In addition, the voice wake-up apparatus and the voice wake-up method provided in the above embodiments belong to the same concept, and the specific implementation process is described in detail in the method embodiments, which will not be repeated here.
[0322] In the above embodiments, all or part can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (for example: coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example: infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium accessible by a computer, or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (for example: floppy disk, hard disk, magnetic tape), optical media (for example: digital versatile disc (DVD)) or semiconductor media (for example: solid state disk (SSD)) and the like. It should be noted that the computer readable storage medium mentioned in the embodiments of the present application can be a non-volatile storage medium, in other words, it can be a non-transitory storage medium.
[0323] It should be understood that "at least one" mentioned herein refers to one or more, and "multiple" refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B; "and / or" herein only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist together, and B exists alone. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, "first", "second" and the like are used to distinguish the same items or similar items with basically the same function and role in the embodiments of the present application. Those skilled in the art can understand that "first", "second" and the like do not limit the number and execution order, and "first", "second" and the like do not necessarily mean different.
[0324] The above describes the embodiments provided by the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method of voice wake-up, the method comprising: The method comprises: voice detection is performed on a bone conduction signal collected by a bone conduction microphone, the bone conduction signal containing command word information input by a sound source; in the case where voice input is detected, an air microphone is turned on, and an air conduction signal is collected by the air microphone; wake-up word detection is performed based on a starting portion of the bone conduction signal and the air conduction signal, the starting portion of the bone conduction signal being determined according to a time difference between a time when the air microphone is turned on and a time when the command word is actually input by the sound source; in the case where the command word is detected to contain the wake-up word, voice wake-up is performed on a device to be woken up.
2. The method of claim 1, wherein, The wake-up word detection based on the starting portion of the bone conduction signal and the air conduction signal comprises: a fusion signal is determined based on the starting portion of the bone conduction signal and the air conduction signal; the fusion signal is subjected to the wake-up word detection.
3. The method of claim 2, wherein, The fusion signal is determined based on the starting portion of the bone conduction signal and the air conduction signal, and comprises: the starting portion of the bone conduction signal and the air conduction signal are fused to obtain the fusion signal; or an enhanced starting signal is generated based on the starting portion of the bone conduction signal, and the enhanced starting signal and the air conduction signal are fused to obtain the fusion signal.
4. The method of claim 2 or 3, wherein, The fusion signal is subjected to the wake-up word detection, and comprises: a plurality of audio frames included in the fusion signal are input into a first acoustic model to obtain a plurality of posterior probability vectors output by the first acoustic model, the plurality of posterior probability vectors corresponding to the plurality of audio frames in a one-to-one manner, and a first posterior probability vector in the plurality of posterior probability vectors is used to indicate a probability that a phoneme of a first audio frame in the plurality of audio frames belongs to a plurality of specified phonemes; the wake-up word detection is performed based on the plurality of posterior probability vectors.
5. The method of claim 1, wherein, The wake-up word detection based on the starting portion of the bone conduction signal and the air conduction signal comprises: a plurality of posterior probability vectors are determined based on the starting portion of the bone conduction signal and the air conduction signal, the plurality of posterior probability vectors corresponding to a plurality of audio frames included in the starting portion of the bone conduction signal and the air conduction signal in a one-to-one manner, and a first posterior probability vector in the plurality of posterior probability vectors is used to indicate a probability that a phoneme of a first audio frame in the plurality of audio frames belongs to a plurality of specified phonemes; the wake-up word detection is performed based on the plurality of posterior probability vectors.
6. The method of claim 5, wherein, The plurality of posterior probability vectors are determined based on the starting portion of the bone conduction signal and the air conduction signal, and comprise: the starting portion of the bone conduction signal and the air conduction signal are input into a second acoustic model to obtain a first number of bone conduction posterior probability vectors and a second number of air conduction posterior probability vectors output by the second acoustic model, the first number of bone conduction posterior probability vectors corresponding to audio frames included in the starting portion of the bone conduction signal in a one-to-one manner, and the second number of air conduction posterior probability vectors corresponding to audio frames included in the air conduction signal in a one-to-one manner; The first bone conduction posterior probability vector and the first air conduction posterior probability vector are fused to obtain the second posterior probability vector. The first bone conduction posterior probability vector corresponds to the last audio frame of the beginning portion of the bone conduction signal, and the duration of the last audio frame is less than the frame duration. The first air conduction posterior probability vector corresponds to the first audio frame of the air conduction signal, and the duration of the first audio frame is less than the frame duration. The plurality of posterior probability vectors include the second posterior probability vector, the vectors other than the first bone conduction posterior probability vector among the first number of bone conduction posterior probability vectors, and the vectors other than the first air conduction posterior probability vector among the second number of air conduction posterior probability vectors.
7. The method of claim 5, wherein, The determination of multiple posterior probability vectors based on the initial portion of the bone conduction signal and the air conduction signal includes: The initial portion of the bone conduction signal and the air conduction signal are input into the third acoustic model to obtain the plurality of posterior probability vectors output by the third acoustic model.
8. The method of any one of claims 4-7, wherein, The detection of the wake word based on the plurality of posterior probability vectors includes: Based on the multiple posterior probability vectors and the phoneme sequence corresponding to the wake word, determine the confidence level of the phoneme sequence corresponding to the command word including the phoneme sequence corresponding to the wake word; If the confidence level exceeds the confidence threshold, it is determined that the command word includes the wake word.
9. The method of any one of claims 4-7, wherein, The detection of the wake word based on the plurality of posterior probability vectors includes: Based on the multiple posterior probability vectors and the phoneme sequence corresponding to the wake word, determine the confidence level of the phoneme sequence corresponding to the command word including the phoneme sequence corresponding to the wake word; If the confidence level exceeds the confidence level threshold and the plurality of posterior probability vectors and the plurality of template vectors satisfy the distance condition, it is determined that the command word includes the wake word, and the plurality of template vectors indicate the probability that the phoneme of the speech signal containing the complete information of the wake word belongs to the plurality of specified phonemes.
10. The method of claim 9, wherein, When the plurality of posterior probability vectors correspond one-to-one with the plurality of template vectors, the distance condition includes: the mean of the distances between the plurality of posterior probability vectors and the corresponding template vectors is less than a distance threshold.
11. The method of claim 9 or 10, wherein, The method further includes: Acquire a bone conduction registration signal, the bone conduction registration signal containing complete information about the wake word; Based on the bone conduction registration signal and the phoneme sequence corresponding to the wake word, the confidence threshold and the plurality of template vectors are determined.
12. The method of claim 11, wherein, The step of determining the confidence threshold and the multiple template vectors based on the bone conduction registration signal and the phoneme sequence corresponding to the wake word includes: The fusion registration signal is determined based on the bone conduction registration signal; Based on the fused registration signal and the phoneme sequence corresponding to the wake word, the confidence threshold and the multiple template vectors are determined.
13. The method of claim 12, wherein, The step of determining the confidence threshold and the multiple template vectors based on the fused registration signal and the phoneme sequence corresponding to the wake-up word includes: inputting the plurality of registration audio frames included in the fusion registration signal into the first acoustic model to obtain a plurality of registration posterior probability vectors output by the first acoustic model, the plurality of registration posterior probability vectors corresponding to the plurality of registration audio frames in a one-to-one manner, and a first registration posterior probability vector in the plurality of registration posterior probability vectors indicating a probability that a phoneme of a first registration audio frame in the plurality of registration audio frames belongs to the plurality of specified phonemes; determining the plurality of template vectors as the plurality of registration posterior probability vectors; determining the confidence threshold based on the plurality of registration posterior probability vectors and the phoneme sequence corresponding to the wake-up word.
14. The method of claim 11, wherein, Before the determining the confidence threshold based on the bone conduction registration signal and the phoneme sequence corresponding to the wake-up word, the method further includes: obtaining an air conduction registration signal; The determining the confidence threshold based on the bone conduction registration signal and the phoneme sequence corresponding to the wake-up word includes: determining a plurality of registration posterior probability vectors based on the bone conduction registration signal and the air conduction registration signal, the plurality of registration posterior probability vectors corresponding to a plurality of registration audio frames included in the bone conduction registration signal and the air conduction registration signal in a one-to-one manner, and a first registration posterior probability vector in the plurality of registration posterior probability vectors indicating a probability that a phoneme of a first registration audio frame in the plurality of registration audio frames belongs to the plurality of specified phonemes; determining the confidence threshold based on the plurality of registration posterior probability vectors and the phoneme sequence corresponding to the wake-up word.
15. An apparatus for voice wake-up, the apparatus comprising: The apparatus includes: a voice detection module configured to perform voice detection based on a bone conduction signal collected by a bone conduction microphone, the bone conduction signal containing command word information input by a sound source; a processing module configured to, in a case where voice input is detected, turn on an air microphone and collect an air conduction signal through the air microphone; a wake-up word detection module configured to, in the case where voice input is detected, perform detection of a wake-up word based on a starting portion of the bone conduction signal and the air conduction signal, the starting portion of the bone conduction signal being determined according to a time difference between a time when the air microphone is turned on and a time when the command word is actually input by the sound source; a voice wake-up module configured to, in a case where the command word is detected to contain the wake-up word, perform voice wake-up on a device to be woken up.
16. The apparatus of claim 15, wherein, The wake-up word detection module includes: a first determination submodule configured to determine a fusion signal based on the starting portion of the bone conduction signal and the air conduction signal; a wake-up word detection submodule configured to perform detection of the wake-up word on the fusion signal.
17. The apparatus of claim 16, wherein, The first determination submodule is configured to: fuse the starting portion of the bone conduction signal and the air conduction signal to obtain the fusion signal; or generate an enhanced starting signal based on the starting portion of the bone conduction signal, and fuse the enhanced starting signal and the air conduction signal to obtain the fusion signal.
18. The apparatus of claim 16 or 17, wherein, The wake-up word detection submodule is configured to: inputting a plurality of audio frames included in the fusion signal into a first acoustic model to obtain a plurality of posterior probability vectors output by the first acoustic model, the plurality of posterior probability vectors corresponding to the plurality of audio frames one by one, a first posterior probability vector in the plurality of posterior probability vectors being used to indicate a probability that a phoneme of a first audio frame in the plurality of audio frames belongs to a plurality of specified phonemes; performing detection of the wake-up word based on the plurality of posterior probability vectors.
19. The apparatus of claim 15, wherein, The wake-up word detection module comprises: a second determination submodule configured to determine a plurality of posterior probability vectors based on the starting portion of the bone conduction signal and the air conduction signal, the plurality of posterior probability vectors corresponding to a plurality of audio frames included in the starting portion of the bone conduction signal and the air conduction signal one by one, a first posterior probability vector in the plurality of posterior probability vectors being used to indicate a probability that a phoneme of a first audio frame in the plurality of audio frames belongs to a plurality of specified phonemes; a wake-up word detection submodule configured to perform detection of the wake-up word based on the plurality of posterior probability vectors.
20. The apparatus of claim 19, wherein, The second determination submodule is configured to: input the starting portion of the bone conduction signal and the air conduction signal into a second acoustic model to obtain a first number of bone conduction posterior probability vectors and a second number of air conduction posterior probability vectors output by the second acoustic model, the first number of bone conduction posterior probability vectors corresponding to audio frames included in the starting portion of the bone conduction signal one by one, the second number of air conduction posterior probability vectors corresponding to audio frames included in the air conduction signal one by one; fuse a first bone conduction posterior probability vector and a first air conduction posterior probability vector to obtain a second posterior probability vector, the first bone conduction posterior probability vector corresponding to a last audio frame of the starting portion of the bone conduction signal, a time length of the last audio frame being less than a frame time length, the first air conduction posterior probability vector corresponding to a first audio frame of the air conduction signal, a time length of the first audio frame being less than the frame time length, the plurality of posterior probability vectors comprising the second posterior probability vector, vectors other than the first bone conduction posterior probability vector in the first number of bone conduction posterior probability vectors, and vectors other than the first air conduction posterior probability vector in the second number of air conduction posterior probability vectors.
21. The apparatus of claim 19, wherein, The second determination submodule is configured to: input the starting portion of the bone conduction signal and the air conduction signal into a third acoustic model to obtain the plurality of posterior probability vectors output by the third acoustic model.
22. The apparatus of any one of claims 18-21, wherein, The wake-up word detection submodule is configured to: determine a confidence degree that a phoneme sequence corresponding to the command word comprises a phoneme sequence corresponding to the wake-up word based on the plurality of posterior probability vectors and the phoneme sequence corresponding to the wake-up word; determine that the command word comprising the wake-up word is detected in a case where the confidence degree exceeds a confidence degree threshold.
23. The apparatus of any one of claims 18-21, wherein, The wake-up word detection submodule is configured to: determine a confidence degree that a phoneme sequence corresponding to the command word comprises a phoneme sequence corresponding to the wake-up word based on the plurality of posterior probability vectors and the phoneme sequence corresponding to the wake-up word; In a case where the confidence exceeds a confidence threshold and the plurality of posterior probability vectors and a plurality of template vectors satisfy a distance condition, the plurality of template vectors indicating probabilities that phonemes of a voice signal containing complete information of the wake-up word belong to the plurality of specified phonemes, it is determined that the command word is detected to include the wake-up word.
24. The apparatus of claim 23, wherein, In a case where the plurality of posterior probability vectors and the plurality of template vectors correspond to each other, the distance condition includes that a mean of distances between the plurality of posterior probability vectors and corresponding template vectors is less than a distance threshold.
25. The apparatus of claim 23 or 24, wherein, The apparatus further includes: an obtaining module, configured to obtain a bone conduction registration signal, the bone conduction registration signal containing complete information of the wake-up word; a determining module, configured to determine the confidence threshold and the plurality of template vectors based on the bone conduction registration signal and a phoneme sequence corresponding to the wake-up word.
26. The apparatus of claim 25, wherein, The determining module includes: a third determining sub-module, configured to determine a fusion registration signal based on the bone conduction registration signal; a fourth determining sub-module, configured to determine the confidence threshold and the plurality of template vectors based on the fusion registration signal and the phoneme sequence corresponding to the wake-up word.
27. The apparatus of claim 26, wherein, The fourth determining sub-module is configured to: input a plurality of registration audio frames included in the fusion registration signal into a first acoustic model to obtain a plurality of registration posterior probability vectors output by the first acoustic model, the plurality of registration posterior probability vectors corresponding to the plurality of registration audio frames one by one, a first registration posterior probability vector in the plurality of registration posterior probability vectors indicating a probability that a phoneme of a first registration audio frame in the plurality of registration audio frames belongs to the plurality of specified phonemes; determine the plurality of template vectors based on the plurality of registration posterior probability vectors; determine the confidence threshold based on the plurality of registration posterior probability vectors and the phoneme sequence corresponding to the wake-up word.
28. The apparatus of claim 25, wherein, The apparatus further includes: an obtaining module, configured to obtain an air conduction registration signal; The determining module includes: a fifth determining sub-module, configured to determine a plurality of registration posterior probability vectors based on the bone conduction registration signal and the air conduction registration signal, the plurality of registration posterior probability vectors corresponding to a plurality of registration audio frames included in the bone conduction registration signal and the air conduction registration signal one by one, a first registration posterior probability vector in the plurality of registration posterior probability vectors indicating a probability that a phoneme of a first registration audio frame in the plurality of registration audio frames belongs to the plurality of specified phonemes; a sixth determining sub-module, configured to determine the confidence threshold based on the plurality of registration posterior probability vectors and the phoneme sequence corresponding to the wake-up word.
29. An electronic device, comprising: The electronic device includes a memory and a processor; The memory is configured to store a computer program; The processor is configured to execute the computer program to implement steps of the method in any of claims 1-14.
30. A computer-readable storage medium, characterized in that, The storage medium has stored therein a computer program, which, when executed by a processor, implements steps of the method in any of claims 1-14.
31. A computer program product comprising computer instructions, characterized in that, The computer instructions, when executed by a processor, implement steps of the method in any of claims 1-14.
Citation Information
Patent Citations
Voice waking-up method and system thereof
CN109036412A
Deep learning speech extraction and noise reduction method fusing bone vibration sensor and microphone signals
CN110931031A
Intelligent microphone and signal processing method thereof
CN113259793A