Voice wake-up method, electronic device, and computer readable storage medium
By combining the improved endpoint detection algorithm and the hidden Markov model, the existing speech wake-up method has solved the problem of large error and calculation overhead in signal preprocessing, and a more efficient and accurate speech wake-up function is achieved.
Patent Information
- Application Number
- PCT/CN2023/135907
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-01
- Publication Date
- 2025-06-05
AI Technical Summary
In the signal preprocessing process, especially the endpoint detection algorithm, the existing voice wake-up method has high error and calculation overhead, resulting in high power consumption and insufficient accuracy.
A voice wake-up method is adopted to collect original speech signals, generate pulse density modulation data, decode and perform preprocessing and feature extraction, and use improved endpoint detection algorithms and hidden Markov models for pattern matching, generate recognition results and wake up the external processor.
While reducing the power consumption of electronic devices, it improves the accuracy of voice wake-up, achieving more efficient voice signal processing and recognition.
Smart Images

Figure CN2023135907_05062025_PF_FP_ABST
Abstract
Description
Voice wake-up method, electronic device, and computer-readable storage medium Technical Field
[0001] The present invention relates to the field of acoustic technology, and in particular to a voice wake-up method, an electronic device, and a computer-readable storage medium. Background Art
[0002] The voice wake-up method in related technologies usually uses a micro-electro-mechanical system (MEMS) microphone to collect voice signals, and then transmits them to the main control chip of the central processing unit (CPU) through an audio transmission interface. The main control chip implements the voice wake-up function by running the algorithm. Among them, the MEMS microphone needs the help of an external CPU to process and recognize the voice signal, and the CPU needs to remain in operation for a long time, which has high power consumption. In addition, the MEMS microphone has a single function, which is not conducive to product modularization. Technical issues
[0003] In the signal preprocessing process of voice wake-up methods, endpoint detection algorithms are often used to accurately find the starting and ending points of a speech signal from a noisy speech signal. Among them, the endpoint detection algorithm is a dual-threshold detection algorithm that combines short-time energy and short-time zero-crossing points to distinguish between silent segments, transition segments, speech segments, and end segments of speech. The dual-threshold detection algorithm has large errors and computational overhead. Technical Solutions
[0004] In view of this, embodiments of the present invention provide a voice wake-up method, an electronic device, and a computer-readable storage medium, for improving the accuracy of voice wake-up while reducing the power consumption of the electronic device.
[0005] In one aspect, an embodiment of the present invention provides a voice wake-up method, applied to an electronic device, comprising:
[0006] Collecting the original voice signal input by the user;
[0007] generating pulse density modulation data according to the original speech signal;
[0008] Decoding the pulse density modulation data to generate decoded data;
[0009] Performing preprocessing and feature extraction processing on the decoded data to generate speech features;
[0010] performing pattern matching on the speech features according to the acquired hidden Markov model to generate a recognition result;
[0011] The external processor of the electronic device is woken up according to the recognition result.
[0012] Optionally, the preprocessing and feature extraction processing of the decoded data to generate speech features includes:
[0013] Preprocessing the decoded data by improving the endpoint detection algorithm to generate valid sound segments of the original speech signal;
[0014] Extracting feature information of the valid sound segment through a feature extraction algorithm;
[0015] Vector quantization is performed on the feature information to generate speech features.
[0016] Optionally, preprocessing the decoded data by using an improved endpoint detection algorithm to generate valid segments of the original speech signal includes:
[0017] filtering interference signals in the decoded data to generate filtered data;
[0018] performing pre-emphasis processing on the filtered data to generate pre-emphasis data;
[0019] Performing frame processing on the pre-emphasized data to generate multiple frames of data;
[0020] Performing windowing processing on each frame of data in the multiple frames of data to generate windowed data;
[0021] Extracting valid content from the windowed data based on the improved endpoint detection algorithm;
[0022] The effective content is calculated based on a Mel-frequency cepstral coefficient feature extraction algorithm to generate effective segments of the original speech signal.
[0023] Optionally, the improved endpoint detection algorithm includes the formula ,in, is the short-time energy change rate, is the short-time energy threshold, is an adjustable impact factor.
[0024] Optionally, performing pattern matching on the speech features according to the acquired hidden Markov model to generate a recognition result includes:
[0025] According to the hidden Markov model, a forward algorithm is used to perform pattern matching on the voice features, so as to determine whether the original voice signal input by the user includes a preset command through a set discrimination rule, and generate a recognition result.
[0026] Optionally, the method of performing pattern matching on the speech features using a forward algorithm based on the hidden Markov model to determine whether the original speech signal input by the user includes a preset command according to a set discrimination rule, and generating a recognition result, includes:
[0027] Using vector quantization to convert the two-dimensional speech features into a one-dimensional symbol sequence;
[0028] Enumerate all possible state sequences corresponding to the symbol sequence of the current frame and generate a feature sequence;
[0029] According to the transition probability and emission probability, the probability of the feature frame sequence being generated by each state sequence is obtained;
[0030] Expand the number of states in each state sequence to the number of feature frames, and sum the probabilities of each state sequence as the likelihood probability of the feature frame sequence being recognized as a word sequence;
[0031] Calculate the probability of the word sequence in the speech model in the feature frame sequence as the prior probability of the word sequence;
[0032] Multiplying the likelihood probability and the prior probability as the posterior probability of the word sequence;
[0033] The word sequence with the largest posterior probability is taken as the recognition result.
[0034] On the other hand, an embodiment of the present invention provides an electronic device, comprising: an external processor and a smart microphone, wherein the smart microphone comprises a microphone and a digital signal processor;
[0035] The microphone is used to collect the original voice signal input by the user;
[0036] The digital signal processor is configured to generate pulse density modulation data based on the original speech signal; decode the pulse density modulation data to generate decoded data; perform preprocessing and feature extraction on the decoded data to generate speech features; perform pattern matching on the speech features based on the acquired hidden Markov model to generate a recognition result; and send the recognition result to the external processor;
[0037] The external processor is configured to perform self-wake-up according to the recognition result.
[0038] Optionally, the digital signal processor receives a hidden Markov model sent by a server, where the hidden Markov model is trained by the server.
[0039] Optionally, the external processor includes a central processing unit or a system-on-chip.
[0040] On the other hand, an embodiment of the present invention provides a computer-readable storage medium, which includes a stored program, wherein when the program is run, the device where the computer-readable storage medium is located is controlled to execute the above-mentioned voice wake-up method. Beneficial effects
[0041] In the technical solution of the voice wake-up method provided by the embodiment of the present invention, the original voice signal input by the user is collected; pulse density modulation data is generated based on the original voice signal; the pulse density modulation data is decoded to generate decoded data; the decoded data is preprocessed and feature extracted to generate voice features; the voice features are pattern matched according to the acquired hidden Markov model to generate recognition results; and the external processor of the electronic device is woken up according to the recognition results. In the technical solution provided by the embodiment of the present invention, the voice wake-up function is realized by preprocessing and feature extracting the decoded data to generate voice features; pattern matching is performed on the voice features according to the acquired hidden Markov model to generate recognition results and the external processor of the electronic device is woken up according to the recognition results, thereby reducing the power consumption of the electronic device while improving the accuracy of voice wake-up. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0043] FIG1 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention;
[0044] FIG2 is a schematic diagram of a circuit structure of an electronic device provided by an embodiment of the present invention;
[0045] FIG3 is a flow chart of a voice wake-up method provided by an embodiment of the present invention;
[0046] FIG4 is a flow chart of performing preprocessing and feature extraction on decoded data to generate speech features according to an embodiment of the present invention;
[0047] FIG5 is a flow chart of a server training a hidden Markov model according to an embodiment of the present invention;
[0048] FIG6 is a flow chart of performing pattern matching on speech features based on the acquired hidden Markov model to generate recognition results. Best Mode for Carrying Out the Invention
[0049] In order to better understand the technical solution of the present invention, the embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0050] It should be understood that the embodiments described are only a portion of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by persons of ordinary skill in the art without creative work are within the scope of protection of the present invention.
[0051] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "an", "the" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.
[0052] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. Furthermore, the character " / " in this document generally indicates an "or" relationship between the associated objects.
[0053] A MEMS microphone is a microphone manufactured using MEMS technology. Simply put, it's a capacitor integrated on a microsilicon wafer. It can be manufactured using a surface-mount process, withstands high reflow soldering temperatures, and is easily integrated with complementary metal oxide semiconductor (CMOS) processes and other audio circuits. It also offers improved noise cancellation and excellent radio frequency (RF) and electromagnetic interference (EMI) suppression. MEMS microphones are often used in products equipped with voice wake-up methods as voice signal acquisition units.
[0054] An embodiment of the present invention provides an electronic device, which includes a true wireless stereo (TWS) headset, a smart watch, a smart light, a smart robot, etc.
[0055] Figure 1 is a schematic diagram of the structure of an electronic device provided by one embodiment of the present invention. As shown in Figure 1, the electronic device includes an external processor 1 and an intelligent microphone 2. The intelligent microphone 2 includes a microphone (MIC) 21 and a digital signal processor (DSP) 22. Microphone 21 is connected to DSP 22, and the external processor 1 is connected to the intelligent microphone 2. Specifically, DSP 22 is connected to the external processor 1.
[0056] The microphone 21 is used to collect the original voice signal input by the user.
[0057] The digital signal processor 22 is used to generate pulse density modulation data based on the original speech signal; decode the pulse density modulation data to generate decoded data; preprocess and extract features from the decoded data to generate speech features; perform pattern matching on the speech features based on the acquired Hidden Markov Model (HMM) to generate recognition results; and send the recognition results to the external processor 1.
[0058] The external processor 1 is used to perform self-wake-up according to the recognition result.
[0059] In one embodiment of the present invention, the external processor 1 includes a central processing unit (CPU) or a system on chip (SOC).
[0060] In one embodiment of the present invention, microphone 21 includes a data output interface and a clock interface, and digital signal processor 22 includes a data input interface, a clock interface, a channel select interface, a reset (RST) interface, a serial clock (SCL) interface, a serial data (SDA) interface, an INT interface, and a wakeup (WAKE) interface. The data output interface of microphone 21 is connected to the data input interface of digital signal processor 22, and the clock interface of microphone 21 is connected to the channel select interface of digital signal processor 22.
[0061] In one embodiment of the present invention, the digital signal processor 22 is specifically used to pre-process the decoded data through an improved endpoint detection algorithm to generate valid sound segments of the original speech signal; extract feature information of the valid sound segments through a feature extraction algorithm; and vector quantize the feature information to generate speech features.
[0062] In one embodiment of the present invention, the digital signal processor 22 is specifically used to filter interference signals in the decoded data to generate filtered data; perform pre-emphasis processing on the filtered data to generate pre-emphasized data; perform frame processing on the pre-emphasized data to generate multi-frame data; perform windowing processing on each frame of data in the multi-frame data to generate windowed data; extract valid content from the windowed data based on an improved endpoint detection algorithm; calculate the valid content based on the Mel-frequency cepstral coefficient feature extraction algorithm to generate a valid sound segment of the original speech signal.
[0063] In one embodiment of the present invention, the improved endpoint detection algorithm includes the formula in, is the short-time energy change rate, is the short-time energy threshold, is an adjustable impact factor.
[0064] In one embodiment of the present invention, the digital signal processor 22 is specifically used to perform pattern matching on speech features using a forward algorithm based on a hidden Markov model, so as to determine whether the original speech signal input by the user includes a preset command through a set discrimination rule, and generate a recognition result.
[0065] In one embodiment of the present invention, the digital signal processor 22 is specifically used to use vector quantization to convert two-dimensional speech features into a one-dimensional symbol sequence; exhaustively enumerate all possible state sequences corresponding to the symbol sequence of the current frame to generate a feature sequence; obtain the probability that the feature frame sequence is generated by each state sequence based on the transition probability and the emission probability; expand the number of states in each state sequence to the number of feature frames, and sum the probabilities of each state sequence as the likelihood probability that the feature frame sequence is recognized as a word sequence; calculate the probability of the word sequence in the feature frame sequence in the speech model as the prior probability of the word sequence; multiply the likelihood probability and the prior probability as the posterior probability of the word sequence; and take the word sequence with the largest posterior probability as the recognition result.
[0066] In one embodiment of the present invention, the digital signal processor 22 receives a hidden Markov model sent by the server, where the hidden Markov model is trained by the server.
[0067] In one embodiment of the present invention, the internal components of the electronic device and its interface connection with the external processor 1 are shown in Figure 1. Microphone 21 samples and modulates the captured voice signal, then outputs the resulting 0-1 digital string via a PDM (Pulse Density Modulation) interface to the digital signal processor 22 for processing and recognition. The recognition result ultimately determines whether to wake up the external processor 1. Furthermore, the external processor 1 can control the digital signal processor 22 via the I2C interface.
[0068] In one embodiment of the present invention, the electronic device has a relatively small size, measuring 3.5mm by 2.65mm. It features high signal-to-noise ratio (SNR), high sensitivity, and low power consumption. It supports keyword recognition, voiceprint recognition, offline recording, and speech recognition, as well as customized wake-up and command words. It can operate at a 3.3V supply voltage and achieve a signal-to-noise ratio (SNR) of 65dB.
[0069] Figure 2 is a schematic diagram of the circuit structure of an electronic device provided by an embodiment of the present invention. As shown in Figure 2, the circuit structure of the electronic device includes: a MIC_OUT interface, a MIC_I interface, a MIC_VDD interface, a MIC_BIAS interface, a VCC interface, a MIC_N interface, a GND interface, a DCDC interface, a VANA interface, an I2C_SDA interface, an I2C_CLK interface, an NRST interface, an INT_0 interface, and a WAKE interface. The MIC_OUT interface is connected to the MIC_I interface, the MIC_I interface is connected to the MIC_VDD interface, the MIC_VDD interface is connected to the MIC_BIAS interface, the MIC_BIAS interface is connected to the VCC interface, the VCC interface is connected to the MIC_N interface, the MIC_N interface is connected to the GND interface, the GND interface is connected to the WAKE interface, the WAKE interface is connected to the INT_0 interface, the INT_0 interface is connected to the NRST interface, the NRST interface is connected to the I2C_CLK interface, the I2C_CLK interface is connected to the I2C_SDA interface, the I2C_SDA interface is connected to the VANA interface, the VANA interface is connected to the DCDC interface, and the DCDC interface is connected to the MIC_OUT interface.
[0070] The MIC_OUT interface is connected to the MIC_I interface through a first capacitor, which can be 1uf; the MIC_VDD interface is connected to the MIC_BIAS interface; the VCC interface is connected to the external power supply; the MIC_N interface is grounded through a second capacitor, which can be 2.2uf; the GND interface is grounded; the I2C_SDA interface, I2C_CLK interface, NRST interface, INT_0 interface and WAKE interface are all connected to the external processor; the VANA interface is grounded through a third capacitor, which can be 2.2nf; the DCDC interface is grounded through a fourth capacitor, which can be 2.2nf.
[0071] Based on the circuit structure of the electronic device in FIG1 or the electronic device in FIG2 , an embodiment of the present invention provides a voice wake-up method. FIG3 is a flow chart of a voice wake-up method provided by an embodiment of the present invention. As shown in FIG3 , the method includes:
[0072] Step 102: Collect the original voice signal input by the user.
[0073] In one embodiment of the present invention, each step is executed by an electronic device.
[0074] In this step, the user inputs an original voice signal to the electronic device by speaking to the electronic device (saying a command word). At this time, the electronic device can collect the original voice signal input by the user.
[0075] Step 104: Generate pulse density modulation data according to the original speech signal.
[0076] In this step, the original speech signal is converted into pulse density modulation (PDM) data in a pulse density modulation (PDM) format.
[0077] Step 106: Decode the pulse density modulation data to generate decoded data.
[0078] Step 108: Preprocess and extract features from the decoded data to generate speech features.
[0079] FIG4 is a flow chart of performing preprocessing and feature extraction on decoded data to generate speech features according to an embodiment of the present invention. As shown in FIG4 , step 108 includes:
[0080] Step 1082: Filter the interference signal in the decoded data to generate filtered data.
[0081] In one embodiment of the present invention, a neural network library based on deep learning and a DSP library for hardware floating-point operations may be used.
[0082] Step 1084: Perform pre-emphasis processing on the filtered data to generate pre-emphasis data.
[0083] In this step, in order to emphasize the high frequency part of the audio, the input filtered data is pre-emphasized.
[0084] Step 1086: perform frame processing on the pre-emphasized data to generate multiple frames of data.
[0085] In this step, in order to meet the Fourier transform requirement for input signal stability and considering that the speech signal is a quasi-steady-state process in a short time range, the pre-emphasized data is framed and the frames are required to overlap.
[0086] Step 1088: Perform windowing processing on each frame of the multiple frames of data to generate windowed data.
[0087] In this step, in order to emphasize the speech waveform of the sample n appendix and eliminate the overlapping parts between frames, a Hamming window is selected to perform windowing processing on each frame of data.
[0088] Step 1090: Extract valid content from the windowed data based on the improved endpoint detection algorithm.
[0089] In this step, in order to distinguish between speech and non-speech areas, an improved endpoint detection algorithm based on a threshold may be used to extract valid content.
[0090] Specifically, the improved endpoint detection algorithm includes the formula ,in, is the short-time energy change rate, is the short-time energy threshold, is an adjustable impact factor.
[0091] The above-mentioned improved endpoint detection algorithm adds a time domain parameter that reflects the degree of change of the signal - the short-time energy change rate. Based on the preset energy threshold initial value, when the change rate of the energy of adjacent signal frames is small, the energy threshold of adjacent frames is also set to a smaller difference, which can reduce the probability of valid signals being incorrectly screened out. In essence, it converts the process of setting a new short-time energy threshold for the signal each time for multiple signal screening into a process of setting a different short-time energy threshold for each frame of the signal for one screening, and finally combines the short-time zero crossing point for secondary screening, thereby optimizing the computational overhead of the algorithm.
[0092] Step 1092: Calculate the effective content based on the Mel-frequency cepstral coefficient feature extraction algorithm to generate effective segments of the original speech signal.
[0093] In this step, the Mel Frequency Cepstrum Coefficient (MFCC) feature extraction algorithm uses fast Fourier transform to extract the corresponding spectrum, and uses a Mel filter bank to reduce the amount of data, mimicking the human ear's high resolution at low frequencies. Finally, the static features of the cepstrum parameters and their corresponding differential spectra are used to improve recognition performance.
[0094] Step 1094: Extract feature information of the valid sound segment through a feature extraction algorithm.
[0095] Step 1096: Perform vector quantization on the feature information to generate speech features.
[0096] Specifically, the feature information is vector quantized to generate a codebook, and the target code vector, ie, the speech feature, is obtained through global search.
[0097] Step 110: Perform pattern matching on the speech features according to the acquired hidden Markov model to generate a recognition result.
[0098] In one embodiment of the present invention, an electronic device can obtain a hidden Markov model (HMM) from a server. The HMM is trained by the server. The server can train various patterns to generate various HMM-based speech models, ultimately creating a template library that can be ported to the DSP.
[0099] In one embodiment of the present invention, an HMM model framework is established, and then the various model parameters obtained from the acoustic model training are input into the corresponding positions in the framework. When a new embedded project is subsequently created, the HMM machine learning sequence model with the input acoustic model parameters is transplanted. The parameters include the state transition matrix, the output probability array of the symbol, and the rejection recognition probability threshold, etc.
[0100] The Hidden Markov Model (HMM) is a stochastic modeling method that primarily analyzes the short-term characteristics of speech and the transition relationships between these characteristics, ultimately calculating likelihood probabilities to make decisions. Figure 5 is a flow chart of a server training a Hidden Markov Model according to one embodiment of the present invention. As shown in Figure 5, the server training process for the HMM is as follows:
[0101] Step S1: After preprocessing an audio signal, a frame sequence and a corresponding word sequence are obtained, which are used as the input of the model and the expectation-maximization (EM) algorithm is used to iteratively solve the problem.
[0102] Step S2: Refine the word sequence into a triphone sequence.
[0103] Step S3: exhaustively enumerate all possible state sequences of the current three-factor sequence to obtain all possible state sequences after the state sequence dimension is expanded to the feature frame dimension.
[0104] Step S4: Initialize the transfer matrix A, the emission matrix B, the initial state probability matrix π, and initialize the probability to be evenly distributed.
[0105] Step S5: Based on the model parameters of the transfer matrix A, the emission matrix B and the initial state probability matrix π obtained in the previous step, the probability of each state sequence is obtained through a forward or backward algorithm.
[0106] Step S6: Calculate each state sequence to obtain the likelihood function of the current triphone sequence, where the transfer matrix A, the emission matrix B, and the initial state probability matrix π are variables.
[0107] Step S7: Calculate the expectation of the triphone sequence in each state sequence and maximize this expectation to obtain the corresponding updated: transfer matrix A, emission matrix B and initial state probability matrix π (derive and set the derivative to zero).
[0108] Step S8: Repeat steps S4, S5, and S6 until the HMM model converges.
[0109] Specifically, according to the Hidden Markov Model, a forward algorithm is used to perform pattern matching on speech features to determine whether the original speech signal input by the user includes a preset command through a set discrimination rule, and generate a recognition result.
[0110] FIG6 is a flow chart of performing pattern matching on speech features based on the acquired hidden Markov model to generate recognition results. As shown in FIG6 , step 110 includes:
[0111] Step 1102: Use vector quantization (VQ) to convert the two-dimensional speech features into a one-dimensional symbol sequence.
[0112] Step 1104: Enumerate all possible state sequences corresponding to the symbol sequence of the current frame to generate a feature sequence.
[0113] Step 1106: Obtain the probability that the feature frame sequence is generated by each state sequence based on the transition probability and the emission probability.
[0114] Step 1108: Expand the number of states in each state sequence to the number of feature frames, and sum the probabilities of each state sequence as the likelihood probability of the feature frame sequence being recognized as a word sequence.
[0115] In this step, a feature frame sequence has far more states than the corresponding words. After the same word sequence is refined into a state sequence, the number of states in the sequence needs to be expanded to the number of feature frames. The expanded state sequence has multiple possibilities. The probabilities of each possible state sequence are summed up as the likelihood probability that the feature frame sequence is recognized as this word sequence.
[0116] Step 1110: Calculate the probability of the word sequence in the speech model in the feature frame sequence as the prior probability of the word sequence.
[0117] Step 1112: Multiply the likelihood probability and the prior probability to obtain the posterior probability of the word sequence.
[0118] Step 1114: Take the word sequence with the largest posterior probability as the recognition result.
[0119]
[0120] in, is the prior probability of the word sequence, which comes from the language model; is the likelihood probability, which comes from the acoustic model ; Where S represents all state sequence combinations that can obtain the observation sequence Y, and h represents a state sequence in S.
[0121] Step 112: Wake up the external processor of the electronic device according to the recognition result.
[0122] In one embodiment of the present invention, a command word input by a user may be matched, and when the match is successful, a wake-up signal may be output to an external processor, and the external processor may wake itself up in response to the wake-up signal.
[0123] In the technical solution provided by the embodiment of the present invention, the original voice signal input by the user is collected; pulse density modulation data is generated based on the original voice signal; the pulse density modulation data is decoded to generate decoded data; the decoded data is preprocessed and feature extracted to generate voice features; the voice features are pattern matched based on the acquired hidden Markov model to generate recognition results; and the external processor of the electronic device is awakened based on the recognition results. In the technical solution provided by the embodiment of the present invention, the voice wake-up function is realized by preprocessing and feature extracting the decoded data to generate voice features; pattern matching is performed on the voice features based on the acquired hidden Markov model to generate recognition results and the external processor of the electronic device is awakened based on the recognition results, thereby reducing the power consumption of the electronic device while improving the accuracy of voice wake-up.
[0124] In the technical solution provided by the embodiment of the present invention, a smart microphone can be used to store 1 to 20 seconds of pulse code modulation (PCM) data, which is then processed by a high-performance DSP algorithm. The data is then triggered to wake up an external processor through the WAKE interface to implement a voice wake-up function, greatly reducing the standby power consumption of the external processor.
[0125] In the technical solution provided by the embodiment of the present invention, the smart microphone includes a MIC and a high-performance DSP, which can realize voice collection, voice processing and voice recognition wake-up operations. Compared with existing MEMS microphones, it is a complete solution with higher computing resources, which is convenient for users to carry out secondary development.
[0126] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program. When the program is running, the device where the computer-readable storage medium is located is controlled to perform the steps of the embodiment of the above-mentioned voice wake-up method. For a specific description, please refer to the embodiment of the above-mentioned voice wake-up method.
[0127] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A voice wake-up method, characterized in that, applied to an electronic device, includes: Collecting the original voice signal input by the user; Generating pulse density modulation data according to the original voice signal; Decoding the pulse density modulation data to generate decoded data; Performing preprocessing and feature extraction on the decoded data to generate voice features; Performing pattern matching on the voice features according to the obtained hidden Markov model to generate an identification result; Waking up the external processor of the electronic device according to the identification result.
2. The method according to claim 1, characterized in that, The performing preprocessing and feature extraction on the decoded data to generate voice features includes: Preprocessing the decoded data through an improved endpoint detection algorithm to generate a valid segment of the original voice signal; Extracting the feature information of the valid segment through a feature extraction algorithm; Performing vector quantization on the feature information to generate voice features.
3. The method according to claim 2, characterized in that, The preprocessing the decoded data through an improved endpoint detection algorithm to generate a valid segment of the original voice signal includes: Filtering the interference signal in the decoded data to generate filtered data; Performing pre-emphasis processing on the filtered data to generate pre-emphasized data; Performing frame segmentation on the pre-emphasized data to generate multi-frame data; Performing windowing processing on each frame of the multi-frame data to generate windowed data; Extracting the valid content in the windowed data based on the improved endpoint detection algorithm; Calculating the valid content based on the Mel Frequency Cepstral Coefficient feature extraction algorithm to generate a valid segment of the original voice signal.
4. The method according to claim 2 or 3, characterized in that, The improved endpoint detection algorithm includes the formula , where is the short-time energy change rate, is the short-time energy threshold, is an adjustable influence factor.
5. The method according to claim 1, characterized in that, The performing pattern matching on the voice features according to the obtained hidden Markov model to generate an identification result includes: Performing pattern matching on the voice features according to the hidden Markov model using the forward algorithm to determine whether the original voice signal input by the user includes a preset command through a set discrimination rule, and generating an identification result.
6. The method according to claim 5, characterized in that, The performing pattern matching on the voice features according to the hidden Markov model using the forward algorithm to determine whether the original voice signal input by the user includes a preset command through a set discrimination rule, and generating an identification result includes: Converting the two-dimensional voice features into a one-dimensional symbol sequence by using vector quantization; Enumerating all possible state sequences corresponding to the symbol sequence of the current frame to generate a feature sequence; Obtaining the probability that the feature frame sequence is generated by each state sequence according to the transition probability and emission probability; Expanding the number of states in each state sequence to the number of feature frames, and summing the probabilities of each state sequence as the likelihood probability that the feature frame sequence is recognized as a word sequence; Calculating the probability of the word sequence in the voice model in the feature frame sequence as the prior probability of the word sequence; Multiplying the likelihood probability and the prior probability as the posterior probability of the word sequence; The word sequence with the largest posterior probability is used as the recognition result.
7. An electronic device, characterized in that, it includes: an external processor and an intelligent microphone, and the intelligent microphone includes a microphone and a digital signal processor; the microphone is used to collect the original voice signal input by the user; the digital signal processor is used to generate pulse density modulation data according to the original voice signal; decode the pulse density modulation data to generate decoded data; perform preprocessing and feature extraction processing on the decoded data to generate voice features; perform pattern matching on the voice features according to the obtained hidden Markov model to generate a recognition result; send the recognition result to the external processor; the external processor is used to perform self-wake-up according to the recognition result.
8. The electronic device according to claim 7, characterized in that, the digital signal processor receives the hidden Markov model sent by the server, and the hidden Markov model is trained by the server.
9. The electronic device according to claim 7, characterized in that, the external processor includes a central processing unit or a system-on-chip.
10. A computer-readable storage medium, characterized in that, the computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the voice wake-up method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Voice control system, wake-up method thereof, wake-up device, household appliance and coprocessor
CN106157950A
Voice awakening method based on soc chip
CN106601229A
Enhanced wake-up method and device for Internet of Things terminal, and storage medium
CN114944153A
Speech recognition device and speech recognition method
US20050119883A1
Analog voice activity detection
US20170263268A1