Intelligent awakening method and device for voice interaction stuffed toy and medium
By using a dual-microphone array and cochlear envelope map processing, combined with a multi-dimensional scoring mechanism and a hidden Markov model, the robustness problem of speech recognition for plush toys in noisy environments was solved, enabling efficient voice interaction wake-up and decoding on low-computing-power devices.
Patent Information
- Application Number
- CN202511327034.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-09-17
AI Technical Summary
In noisy or multi-speaker environments, single-channel speech recognition struggles to distinguish target speech from interfering speech. Existing methods lack robustness in low-computing-power devices such as plush toys, and conventional methods are prone to losing weak speech components, resulting in compromised speech integrity.
A dual-microphone array is used to collect speech signals. The cochlear envelope map is generated by processing the signal through short-time Fourier transform and bandpass filter bank. The sound source mask is formed by combining the time spectrum and Euclidean distance. The wake word and semantic trigger word are decoded using a hidden Markov model. The sound source is identified by combining density, harmonic and prosodic features.
It achieves accurate extraction of the main sound source and acquisition of pure speech signal under low computing power conditions, improves the robustness and accuracy of voice interaction, avoids dependence on high computing power, and is suitable for low-power devices such as plush toys.
Smart Images

Figure CN120833780A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of voice interaction, and in particular to a smart wake-up method, device and medium for a voice interaction plush toy. BACKGROUND
[0002] With the rapid development of voice interaction technology, single-channel speech recognition has been widely used in smart speakers, vehicle-mounted systems and voice assistants. However, in noisy or multi-speaker environments, single-channel processing often has difficulty in distinguishing target speech from interfering speech. To improve the robustness of voice interaction, researchers have proposed multi-channel signal processing methods, such as beamforming, generalized cross-correlation phase transform (GCC-PHAT) and threshold separation based on power spectrum. These methods use a multi-microphone array to enhance the target speech and reduce the impact of noise, and to some extent, improve the speech recognition effect.
[0003] However, the existing methods generally have two types of deficiencies: first, relying on a single physical acoustics parameter (such as power or cross-correlation time delay) for sound source discrimination, resulting in insufficient robustness in dynamic environments; second, the conventional binary mask method is prone to losing weak speech components, causing damage to the integrity of the speech. In particular, in the context of voice interaction for plush toys with low-power chips, computational power and energy consumption are limited, making it difficult to directly apply computationally complex methods such as deep neural networks. Therefore, how to achieve sound source separation and wake-up recognition close to the human auditory mechanism under limited computing conditions has become a core problem that needs to be solved. SUMMARY
[0004] In view of the above-mentioned existing problems, the present application is proposed.
[0005] Therefore, the present application provides a smart wake-up method for a voice interaction plush toy to solve the problem of accurate wake-up recognition in the voice interaction process of the plush toy.
[0006] To solve the above technical problems, the present application provides the following technical solutions: In a first aspect, the present application provides a smart wake-up method for a voice interaction plush toy, comprising, performing short-time Fourier transform on the collected original speech signal, and processing it using a band-pass filter bank to obtain the time-frequency spectrum and cochlear envelope of the original speech signal, forming a two-dimensional statistical graph based on the amplitude attenuation ratio and time delay of each time-frequency point in the time-frequency spectrum, performing local peak value detection on the two-dimensional statistical graph to obtain peak time-frequency points, and performing classification and distribution based on the Euclidean distance of all time-frequency points to the peak time-frequency points to obtain a candidate sound source mask set; comprehensively ranking the candidate sound source mask set according to the density score, harmonic score and prosodic feature score, and selecting the highest ranked candidate sound source mask as the main sound source mask; Refining and weighting the main sound source mask based on the amplitude attenuation ratio and the time delay, generating a main sound source soft mask, and obtaining a pure speech signal; Using a hidden Markov model of a preset wake-up word and a semantic trigger word, outputting a pure phoneme sequence of the pure speech signal, comparing the pure phoneme sequence with a basic phoneme set, and determining a wake-up state of the plush toy according to a comparison result.
[0007] As a preferred scheme of the intelligent wake-up method of the voice interaction plush toy, the classification and distribution based on the Euclidean distance from all time-frequency points to the peak time-frequency point are specifically as follows, The Euclidean distance from all time-frequency points to the peak time-frequency point is calculated, and the mean and standard deviation of the Euclidean distance are calculated; Based on the mean and standard deviation of the Euclidean distance, a similarity threshold is set; If the Euclidean distance from the time-frequency point to the peak time-frequency point is greater than the similarity threshold, it is determined that the sound source corresponding to the time-frequency point is an interference sound source, otherwise, it is determined that the sound source corresponding to the peak time-frequency point is classified into the candidate sound source to which the peak time-frequency point belongs.
[0008] As a preferred scheme of the intelligent wake-up method of the voice interaction plush toy, the cochlear envelope diagram refers to inputting an original speech signal into a filter bank composed of a plurality of equivalent rectangular bandwidth filters, and outputting a bandpass signal of each equivalent rectangular bandwidth filter; The Hilbert transform is performed on each bandpass signal to obtain a bandpass analytic signal, and the module length of the bandpass analytic signal is taken as an energy envelope value of each bandpass; Taking the bandpass index as the vertical axis and the time frame index as the horizontal axis, the energy envelope values of all bandpasses are combined to form a cochlear envelope diagram.
[0009] As a preferred scheme of the intelligent wake-up method of the voice interaction plush toy, the two-dimensional statistical diagram is formed by the following specific steps, Based on the time-frequency spectrum of the original speech signal, the amplitude attenuation ratio of each time-frequency point and the time delay corresponding to the phase difference are calculated; Each time-frequency point amplitude attenuation ratio and time delay are taken as two-dimensional features and mapped to a two-dimensional plane to form a two-dimensional statistical diagram.
[0010] As a preferred scheme of the intelligent wake-up method of the voice interaction plush toy, the candidate sound source mask set is comprehensively evaluated according to the density score, the harmonic score and the prosodic feature score by the following specific steps, For each candidate sound source mask in the candidate sound source set, the number of time-frequency points is counted and normalized to obtain a density score; performing inverse short-time Fourier transform on the time-frequency point corresponding to the candidate sound source mask to obtain a reconstructed speech signal corresponding to the candidate sound source mask, performing Fourier transform on an autocorrelation sequence of the reconstructed speech signal to obtain a spectral correlation function, and performing weighted summation on fundamental frequencies and harmonic frequencies in the spectral correlation function to obtain a harmonic score; Based on the human speech prosody features and the auditory perception rules, a target modulation frequency band range is set, and a prosody feature score is obtained by calculating the energy ratio of the cochlear envelope in the target modulation frequency band range to the energy in the full modulation frequency band range; The density score, the harmonic score and the prosody feature score are weighted and summed to obtain a comprehensive evaluation value of the candidate sound source mask.
[0011] As a preferred scheme of the intelligent wake-up method of the speech interaction plush toy, the specific steps of outputting the pure phoneme sequence of the pure speech signal are as follows, Frame the pure speech signal, calculate the power spectrum of each frame signal after Hamming window weighting, and output a speech feature vector through a mel filter bank; The wake-up word and the semantic trigger word preset in the plush toy are transcribed into a basic phoneme set by constructing a pinyin phoneme mapping table; For each phoneme in the basic phoneme set, a hidden Markov model is constructed respectively; A training speech feature vector set is obtained by collecting and processing a training speech set including the wake-up word and the semantic trigger word; The hidden Markov model of each basic phoneme is trained based on the training phoneme set and the training speech feature vector set; Based on the speech feature vector of the pure speech signal and the hidden Markov model, the Viterbi algorithm is used to search for an optimal state path; Based on the state relationship of each hidden Markov model and the basic phoneme, the optimal state path is converted into a pure phoneme sequence.
[0012] As a preferred scheme of the intelligent wake-up method of the speech interaction plush toy, the specific steps of comparing the pure phoneme sequence with the basic phoneme set and determining the wake-up state of the plush toy according to the comparison result are as follows, If the pure phoneme sequence only contains the phonemes of the wake-up word, it is determined that the plush toy enters a light wake-up state and prompts, and if the new pure phoneme sequence received within a short time window contains the phonemes of the semantic trigger word, the user is interacted with according to the function corresponding to the semantic trigger word, otherwise, the plush toy enters a standby state; If the pure phoneme sequence contains the phonemes of the wake-up word and the semantic trigger word, the user is directly interacted with according to the function corresponding to the semantic trigger word; If the pure phoneme sequence does not include the phonemes of the wake-up word and the semantic trigger word, no interaction is performed, and a standby state is maintained.
[0013] As a preferred scheme of the intelligent wake-up method of the voice interaction plush toy, the original voice signal is synchronously collected through a double microphone array arranged on the left and right sides of the plush toy. The original voice signal includes a left original voice signal collected by the left microphone and a right original voice signal collected by the right microphone.
[0014] In a second aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and wherein the computer program, when executed by the processor, implements any step of the intelligent wake-up method of the voice interaction plush toy according to the first aspect of the present application.
[0015] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements any step of the intelligent wake-up method of the voice interaction plush toy according to the first aspect of the present application.
[0016] The present application has the following beneficial effects: The present application constructs a time-frequency analysis and a cochlea envelope of a double-channel signal, generates and optimizes a sound source mask by combining amplitude attenuation ratio and time delay consistency, thereby realizing accurate extraction of a main sound source, improves the robustness of main sound source determination through a multi-dimensional scoring mechanism of density, harmonic nature and prosodic features, effectively avoids the limitations of relying on a single physical parameter, realizes accurate decoding of a wake-up word and a semantic trigger word based on a lightweight modeling of a phoneme-level hidden Markov model, and avoids high computing power dependence of a large-scale neural network. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0018] Fig. 1 The flowchart of the intelligent wake-up method of the voice interaction plush toy.
[0019] Fig. 2 The flowchart of candidate sound source mask generation and classification.
[0020] Fig. 3 The flowchart of candidate sound source mask evaluation and main sound source determination.
[0021] Fig. 4 The flowchart of wake-up state determination. DETAILED DESCRIPTION
[0022] In order to make the above objectives, characteristics and advantages of the present application more apparent, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0023] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without the specific details set forth in this description. In other instances, well-known methods, procedures, components, and circuits have not been described in detail as not to unnecessarily obscure aspects of the present application.
[0024] Secondly, the "one embodiment" or "embodiment" referred to herein means a specific feature, structure, or characteristic that can be included in at least one implementation of the present application. The "in one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an embodiment that is independent of or mutually exclusive with other embodiments.
[0025] Reference Figs. 1-4 For one embodiment of the present application, the embodiment provides a smart wake-up method of a voice interactive plush toy, comprising the following steps: S1, respectively, the original speech signal collected is subjected to short-time Fourier transform, and a band-pass filter bank is used for processing to obtain a time-frequency spectrum and a cochlea envelope of the original speech signal, a two-dimensional statistical graph is formed based on the amplitude attenuation ratio and time delay of each time-frequency point in the time-frequency spectrum, local peak value detection is performed on the two-dimensional statistical graph to obtain peak time-frequency points, and classification and distribution are performed based on the Euclidean distance of all time-frequency points to the peak time-frequency points to obtain a candidate sound source mask set.
[0026] Two microphones are arranged on the left and right sides of the plush toy, a double microphone array is formed through the spatial interval, the double microphone array is synchronously collected to obtain the original speech signal, including the left original speech signal collected by the left microphone and the right original speech signal collected by the right microphone, the original speech signal is divided into frames and windowed according to a fixed frame length to obtain the time-domain waveform of each frame, the frequency domain representation of each frame of original speech signal is calculated respectively, the time-frequency spectrum of the original speech signal is obtained through short-time Fourier transform; The time-frequency spectrum is a complex value obtained by mapping the original speech signal after weighting by a window function to the frequency domain, each frequency point of the time-frequency spectrum contains two parts of amplitude and phase, the modulus of the time-frequency spectrum is the amplitude spectrum of the original speech signal, and the angle is the phase spectrum of the original speech signal; The original speech signal is input into a filter bank containing several equivalent rectangular bandwidth filters, and the complex original speech signal is decomposed into several non-overlapping or partially overlapping frequency bands to obtain the bandpass signal of each equivalent rectangular bandwidth filter. The bandpass analysis signal is obtained by performing Hilbert transform on each bandpass signal, and the modulus of the bandpass analysis signal is used as the energy envelope value of each bandpass, which mainly reflects the intensity fluctuation and rhythm change of the speech, that is, the prosodic information. The energy envelope values of all bandpasses are combined to form a cochlear envelope map, where the vertical axis is the bandpass index and the horizontal axis is the time frame index; It should be noted that a single speech signal contains rich frequency components, and different frequency intervals carry different speech feature information. The bandwidth design of the equivalent rectangular bandwidth filter bank follows the frequency resolution law of the human cochlea (high resolution at low frequencies and low resolution at high frequencies). Therefore, the structure of the speech signal after decomposition by the filter bank is similar to the neural signal processing method of the human ear before the auditory cortex. Therefore, the original speech signal is input into a filter bank containing several equivalent rectangular bandwidth filters for processing. In essence, this is to simulate the frequency resolution mechanism of the human cochlea, that is, to perform multi-channel decomposition of the original speech signal similar to that of the basilar membrane of the cochlea. Based on the time-frequency spectrum of the original speech signal, the amplitude attenuation ratio of each time-frequency point and the time delay corresponding to the phase difference are calculated as follows: ; ; Where, Represents the original speech signal in time frame No. The amplitude attenuation ratio of the frequency point, and Respectively represent the left original speech signal and the right original speech signal in the time frame No. The time-frequency spectrum of the frequency point, that is, the instant frequency point, is the frequency point index, is the timeframe index, is the original speech signal in time frame No. The time delay corresponding to the phase difference of the frequency point indicates the time delay caused by the sound wave propagating between the left and right microphones. The left original speech signal and the right original speech signal in time frame No. The phase difference of the frequency point, and Respectively represent the left original speech signal and the right original speech signal in the time frame No. The phase of each frequency point; The amplitude attenuation ratio and the time delay of each time-frequency point are taken as two-dimensional features and mapped to a two-dimensional plane to form a two-dimensional statistical diagram, wherein the horizontal axis of the two-dimensional statistical diagram is the amplitude attenuation ratio, and the vertical axis is the time delay; The two-dimensional statistical diagram has high-density peaks at the real direction of the sound source, and each peak corresponds to a candidate sound source. Based on the use scenario of the plush toy, it is known that there are several sound sources when the dual-microphone array collects voice signals. The peak search method is used to detect local peaks of the two-dimensional statistical diagram to obtain the time-frequency points corresponding to the peaks, i.e., the peak time-frequency points, and the amplitude attenuation ratio and the time delay. Based on the Euclidean distance of all time-frequency points to the peak time-frequency points, a candidate sound source mask set is obtained. Specifically, the Euclidean distance of all time-frequency points to the peak time-frequency points is calculated, and the mean and standard deviation of the Euclidean distance are calculated. Based on the mean and standard deviation of the Euclidean distance, a similarity threshold is set by the standard deviation method. By comparing the Euclidean distance with the similarity threshold, if the Euclidean distance of the time-frequency point to the peak time-frequency point is greater than the similarity threshold, it is determined that the sound source corresponding to the time-frequency point is an interference sound source, otherwise, it is determined that the sound source corresponding to the peak time-frequency point is classified to the candidate sound source to which the peak time-frequency point belongs, and a candidate sound source mask is generated. The calculation formula is as follows: ; In the formula, is the candidate sound source mask, is the Euclidean distance of the time-frequency point to the peak time-frequency point, is the similarity threshold; The candidate sound source masks of all peak time-frequency points are summarized to obtain a candidate sound source mask set. It should be noted that the "sound source" in the candidate sound source mask set does not directly refer to the "position or individual from which sound is emitted" in the real world in the physical sense, but refers to a "time-frequency point set" aggregated by acoustic features at the signal processing level. The candidate sound source mask is automatically determined by peak detection in the two-dimensional statistical diagram, and the entire process relies on the consistency of acoustic features rather than external set conditions. The single-step clustering method driven by peaks avoids the burden of plush toys in terms of computing power and power consumption. At the same time, the Euclidean distance as a measurement standard is intuitive, has small calculation amount and strong real-time performance, and meets the low-power real-time processing requirements of toy devices.
[0027] S2, the candidate sound source mask set is ranked according to the density score, the harmonic score and the prosody feature score, and the candidate sound source mask with the highest ranking is selected as the main sound source mask.
[0028] For each candidate sound source mask in the candidate sound source set, the number of time-frequency points is counted and normalized to obtain a density score. It should be noted that the greater the number of time-frequency points, the higher the concentration of time-frequency points in the candidate sound source mask, indicating that the candidate sound source mask is consistent in different frequency bands of the speech signal, and the stronger the reliability. The time-frequency points corresponding to the candidate sound source mask are subjected to inverse short-time Fourier transform to obtain a reconstructed speech signal corresponding to the candidate sound source mask. The autocorrelation sequence of the reconstructed speech signal is subjected to Fourier transform to obtain a spectral correlation function. The fundamental frequency and harmonic frequency in the spectral correlation function are subjected to weighted summation to obtain a harmonic score. The specific steps are as follows: The time-frequency spectrum of the original speech signal is subjected to delay summation beamforming to obtain a basic time-frequency spectrum of the reconstructed speech signal. Specifically, a microphone that receives the speech signal earlier is selected as a phase reference point. According to the time delay generated by the propagation of the left and right microphones, an exponential compensation factor is calculated using a frequency domain phase compensation method to compensate the phase of the speech signal of the other microphone, so that the speech signals of the left and right microphones are aligned. The calculation formula is as follows: ; In the formula, is the exponential compensation factor, is the imaginary unit, is the circular constant, is the exponential base; The time-frequency spectrum of the original speech signal is weighted and summed based on the exponential compensation factor to obtain a basic time-frequency spectrum. Based on the basic time-frequency spectrum and the candidate sound source mask, a candidate time-frequency spectrum corresponding to the candidate sound source mask is calculated. The candidate time-frequency spectrum is subjected to inverse short-time Fourier transform, and each time frame is added to obtain a reconstructed speech signal. The calculation formula of the candidate time-frequency spectrum is as follows: ; In the formula, is the candidate time-frequency spectrum, is the basic time-frequency spectrum; Based on the reconstructed speech signal, the autocorrelation method is used to calculate the autocorrelation sequence of the speech signal. The calculation formula is as follows: ; In the formula, represents the autocorrelation sequence of the speech signal, and are the amplitudes of the reconstructed speech signal of the time frames and , is the delay, is the frame length of the reconstructed speech signal; The spectral correlation function is obtained by Fourier transforming the autocorrelation sequence of the speech signal. Harmonic peak detection is performed on the spectrum correlation function to obtain a fundamental frequency and harmonic frequencies of the reconstructed speech signal, wherein the harmonic frequencies are integer multiples of the fundamental frequency; The fundamental frequency and the first several harmonic frequencies of the reconstructed speech signal are weighted and summed, and normalized to obtain a harmonic score; The energy envelope values of each bandpass of the cochlear envelope are normalized by sliding mean to obtain normalized envelopes of each bandpass, and the normalized envelopes of each bandpass are linearly aggregated with equal weights to obtain a full-band envelope; The full-band envelope is de-trended by sliding mean to eliminate the slow drift trend of the full-band envelope and highlight the rhythm fluctuations of the original speech signal to obtain a de-trended envelope, and the de-trended envelope is subjected to a Hanning window and a discrete Fourier transform to obtain a modulation spectrum; Based on the human speech rhythm characteristics and the auditory perception rules, a target modulation frequency band range, for example, 4Hz~16Hz, is set, the modulation spectrum is subjected to bandpass filtering integration in the target modulation frequency band range and the full modulation frequency band range respectively to obtain a target band energy and a full-band energy, and a rhythm characteristic score is defined as the ratio of the target band energy to the full-band energy; The density score, the harmonic score and the rhythm characteristic score are weighted and summed to obtain a comprehensive evaluation value of the candidate sound source mask, and the candidate sound source mask with the highest comprehensive evaluation value is taken as the main sound source mask; It should be noted that the conventional method for selecting the candidate sound source mask, such as the power threshold determination method, the cross-correlation time delay estimation method and the generalized cross-correlation phase transform, all rely on a single physical acoustics parameter, while the three types of score fusion mechanisms are used to distinguish the candidate sound source mask, which not only considers the spatial statistical characteristics, but also introduces the physiological and auditory features of the speech itself, realizes the sound source selection closer to the human auditory mechanism, and makes the plush toy have stronger anti-interference ability in speech interaction.
[0029] S3, based on the amplitude attenuation ratio and the time delay, the main sound source mask is refined and weighted to generate a main sound source soft mask, and a pure speech signal is obtained.
[0030] Based on the amplitude attenuation ratio and the time delay of the peak time-frequency point and the remaining time-frequency points in the main sound source mask, a Gaussian weighting method is used to calculate the consistency weight of the time-frequency point, and the main sound source mask is refined and weighted by using the consistency weight, wherein the consistency weight of the peak time-frequency point is 1, to obtain a main sound source soft mask, wherein the calculation formula of the consistency weight is as follows: ; In the formula, is the consistency weight of the time-frequency point , and and respectively represent the amplitude attenuation ratio and the time delay of the peak time-frequency point , denotes a peak time-frequency point, and respectively denote the scale parameters of the amplitude attenuation ratio and the time delay, which are obtained by taking the median absolute deviation of the amplitude attenuation ratio and the time delay of all time-frequency points in the main sound source mask respectively; The main sound source soft mask is multiplied by the basic time-frequency spectrum element by element to obtain a soft masking time-frequency spectrum, and the soft masking time-frequency spectrum is subjected to short-time Fourier transform and overlap addition to obtain a pure speech signal. It should be noted that for the voice interaction scene of the plush toy, the traditional binary mask is easy to misdelete weak voice, causing wake-up failure, and the use of amplitude attenuation ratio and time delay to calculate the consistency weight of the time-frequency point makes the retention degree of each time-frequency point not a hard two choices, but a continuous weight distribution, which is more in line with the auditory characteristics of the human ear, and can balance the voice integrity and noise suppression; S4, using the preset wake-up word and semantic trigger word hidden Markov model, output the pure phoneme sequence of the pure speech signal, compare the pure phoneme sequence with the basic phoneme set, and determine the wake-up state of the plush toy according to the comparison result.
[0031] The pure speech signal is framed, the power spectrum of each frame signal is calculated after Hamming window weighting, and the speech feature vector is output through the mel filter bank. Specifically, the pure speech signal is framed, the power spectrum is calculated after Hamming window weighting of each frame of pure speech signal, and the power spectrum is input into the mel filter bank. The mel filter bank is composed of a plurality of mel bandpass filters, the filter center frequencies are uniformly distributed according to the mel scale, the output of each mel bandpass filter is weighted and summed to obtain the mel energy, the mel energy of each frame of pure speech signal is logarithmically transformed, and the discrete cosine transform is performed to obtain the mel cepstrum coefficient of each frame of pure speech signal. The mel cepstrum coefficients of all frames are spliced into a vector to output the speech feature vector; According to the pinyin rule, the pinyin of the preset wake-up word and semantic trigger word (for example, the wake-up word: Mao Mao, the semantic trigger word: play, speak and sing, etc.) of the plush toy is divided into initial and final and tone, the wake-up word and semantic trigger word are transcribed into a basic phoneme set by constructing a pinyin phoneme mapping table, a training speech set including the wake-up word and semantic trigger word and a corresponding text information set are obtained, each piece of text information in the text information set is segmented, each subword is converted into pinyin, and the pinyin phoneme mapping table is used to transcribe the training phoneme set. Align each training speech in the training speech set with the corresponding training phoneme to obtain the boundary information of each phoneme in the time frame of each training speech, and label each training speech to obtain a labeled training speech set. Each labeled training speech is framed, the power spectrum of each frame signal is calculated after Hamming window weighting, and the training speech feature vector set is output through the mel filter bank. For each phoneme in the basic phoneme set, a hidden Markov model is constructed respectively, wherein the hidden Markov model is composed of a plurality of states, respectively representing the starting stage, stable stage and ending stage of each phoneme in the pronunciation process, the maximum likelihood estimation method is used to calculate the state transition probability between each state, for describing the possibility of the phoneme transition from one stage to the next stage in the time dimension, and the observation probability distribution of each state is calculated based on the Gaussian distribution, for describing the probability of the speech feature vector appearing in each state; The hidden Markov model of each basic phoneme is trained based on the training phoneme set and the training speech feature vector set, wherein the model parameters of the hidden Markov model are iteratively estimated by the Baum-Welch algorithm (forward-backward algorithm), and the state transition probability and the observation probability distribution are iteratively updated, and when the maximum iteration number is reached, the training is ended; Based on the speech feature vector of the pure speech signal and the hidden Markov model of the basic phoneme set, the maximum probability value of each state at each time is calculated by using the Viterbi algorithm in a frame-by-frame recursive manner, and the optimal transition source is recorded to obtain the optimal state path, specifically, at the starting time, the first frame of the speech feature vector and the state transition probability between each state in the hidden Markov model of each basic phoneme are determined according to the initial state probability, then starting from the second frame, the state transition probability from all possible states of the previous frame to the current state is calculated on each frame, and the state transition probability is combined with the observation probability distribution of the current state, the path with the maximum state transition probability is selected as the candidate of the next state path, and the predecessor state index of the current state is recorded, finally the maximum state transition probability of each state is obtained on the last frame of the speech feature vector, and the globally optimal termination state is selected from them, and the optimal state path is obtained by tracing back along the recorded predecessor state index from the termination state; Further, based on the correspondence between the states of each hidden Markov model and the basic phonemes, the optimal state path is converted into a pure phoneme sequence; It should be noted that the current method for decoding phonemes mainly uses deep neural networks or models based on words and syllables, both of which require a large amount of speech data for model training, and in the present application, the phoneme is taken as the smallest modeling unit, the hidden Markov model is used to describe the time evolution characteristics of the phoneme, and only the wake-up word and the semantic trigger word are needed to complete the decoding, which is more suitable for the case of low chip power consumption, small computing power and limited data storage space of the plush toy; The pure phoneme sequence is compared with the basic phoneme set, and according to the comparison result, the wake-up state of the plush toy is determined. Specifically, if the pure phoneme sequence only contains the phonemes of the wake-up word, it is determined that the plush toy enters a light wake-up state, and the full-power reception is maintained within a short time window (such as 5 seconds) and a prompt is performed, such as lighting and playing a prompt sound. If the new pure phoneme sequence received within the short time window contains the phonemes of the semantic trigger word, the user is interacted with voice according to the function corresponding to the semantic trigger word, otherwise, the plush toy enters a standby state. If the pure phoneme sequence contains the phonemes of the wake-up word and the semantic trigger word, the user is directly interacted with voice according to the function corresponding to the semantic trigger word. If the pure phoneme sequence does not contain the phonemes of the wake-up word and the semantic trigger word, no interaction is performed, and the standby state is maintained.
[0032] The embodiment also provides a computer device suitable for the smart wake-up method of the voice interactive plush toy, which includes a memory and a processor. The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to realize the smart wake-up method of the voice interactive plush toy proposed in the above embodiment.
[0033] The computer device can be a terminal, which includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, operator network, NFC (near field communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device. In addition, the input device can be an external keyboard, touchpad or mouse, etc.
[0034] The embodiment also provides a storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the intelligent wake-up method of the voice interaction plush toy proposed in the above embodiment; and the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disk.
[0035] To sum up, the application realizes the accurate extraction of the main sound source by constructing the time-frequency analysis and cochlea envelope of the dual-channel signal, combining the amplitude attenuation ratio and time delay consistency to generate and optimize the sound source mask, improves the robustness of the main sound source determination through the multi-dimensional scoring mechanism of the density, harmonic and prosody characteristics, effectively avoids the limitation of relying on a single physical parameter, realizes the accurate decoding of the wake-up word and semantic trigger word based on the lightweight modeling of the phoneme-level hidden Markov model, and avoids the high computing power dependence of a large-scale neural network.
[0036] It should be noted that the above embodiments are only used to illustrate the technical solutions of the application rather than limit the application. Although the application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the application, and all of them should be covered in the scope of the claims of the application.
Claims
1. A method for smart wake-up of a voice interactive plush toy, characterized in that: include, The collected original speech signal is subjected to short-time Fourier transform and processed using a bandpass filter bank to obtain the time-frequency spectrum and cochlear envelope of the original speech signal. A two-dimensional statistical graph is formed based on the amplitude attenuation ratio and time delay of each time-frequency point in the time-frequency spectrum. Local peak detection is performed on the two-dimensional statistical graph to obtain the peak time-frequency point. All time-frequency points are classified and allocated based on the Euclidean distance from the peak time-frequency point to obtain a set of candidate sound source masks. The candidate sound source mask set is comprehensively ranked according to the density score, harmonicity score and rhythmic feature score, and the candidate sound source mask with the highest ranking is selected as the main sound source mask; The main sound source mask is refined and weighted based on the amplitude attenuation ratio and time delay to generate a soft mask of the main sound source and obtain a pure speech signal; Using the hidden Markov model of preset wake-up words and semantic trigger words, the pure phoneme sequence of the pure speech signal is output, and the pure phoneme sequence is compared with the basic phoneme set. According to the comparison results, the awakening state of the plush toy is determined.
2. The intelligent awakening method for a voice-interactive plush toy according to claim 1, characterized in that: The classification and allocation is performed based on the Euclidean distance from all time-frequency points to the peak time-frequency point. The specific steps are as follows: Calculate the Euclidean distance from all time-frequency points to the peak time-frequency point, and calculate the mean and standard deviation of the Euclidean distance; Set the similarity threshold based on the mean and standard deviation of the Euclidean distance; If the Euclidean distance from the time-frequency point to the peak time-frequency point is greater than the similarity threshold, the sound source corresponding to the time-frequency point is determined to be an interference sound source. Otherwise, the sound source corresponding to the peak time-frequency point is classified as the candidate sound source to which the peak time-frequency point belongs. 3.The method of claim 1, wherein: The cochlear envelope diagram refers to inputting the original speech signal into a filter bank composed of several equivalent rectangular bandwidth filters and outputting the bandpass signals of each equivalent rectangular bandwidth filter; Perform Hilbert transform on each bandpass signal to obtain a bandpass analytical signal, and use the modulus of the bandpass analytical signal as the energy envelope value of each bandpass; With the bandpass index as the vertical axis and the time frame index as the horizontal axis, the energy envelope values of all bandpasses are combined to form the cochlear envelope map.
4. The method of smart wake-up of a voice interactive plush toy as claimed in claim 1, wherein: The specific steps of forming a two-dimensional statistical graph are as follows: Based on the time-frequency spectrum of the original speech signal, calculate the amplitude attenuation ratio of each time-frequency point and the time delay corresponding to the phase difference; The amplitude attenuation ratio and time delay of each time-frequency point are used as two-dimensional features and mapped to a two-dimensional plane to form a two-dimensional statistical graph.
5. The method for smart wake-up of a voice interactive plush toy according to claim 1, wherein: The candidate sound source mask set is comprehensively evaluated according to the density score, harmonicity score and rhythmic feature score. The specific steps are as follows: For each candidate sound source mask in the candidate sound source set, count the number of time-frequency points and normalize them to obtain a density score; Perform an inverse short-time Fourier transform on the time-frequency points corresponding to the candidate sound source mask to obtain the reconstructed speech signal corresponding to the candidate sound source mask, perform a Fourier transform on the autocorrelation sequence of the reconstructed speech signal to obtain a spectral correlation function, perform a weighted summation on the fundamental frequency and harmonic frequencies in the spectral correlation function to obtain a harmonicity score; Based on the prosodic characteristics of human speech and the laws of auditory perception, a target modulation frequency band is set, and the prosodic feature score is obtained by calculating the energy ratio of the cochlear envelope within the target modulation frequency band and the full modulation frequency band. The density score, the harmonic score and the prosody feature score are weighted and summed to obtain a comprehensive evaluation value of the candidate sound source mask.
6. The method of smart wake-up of a voice interactive plush toy as claimed in claim 1, wherein: The specific steps are as follows, The clean speech signal is framed, and the power spectrum of each frame signal is calculated after being weighted by a Hamming window, and a speech feature vector is output through a Mel filter bank; The preset wake-up word and semantic trigger word of the plush toy are transcribed into a basic phoneme set by constructing a pinyin phoneme mapping table; For each phoneme in the basic phoneme set, a hidden Markov model is constructed; A training speech feature vector set is obtained by collecting and processing a training speech set including a wake-up word and a semantic trigger word; The hidden Markov model of each basic phoneme is trained based on the training phoneme set and the training speech feature vector set; Based on the speech feature vector of the clean speech signal and the hidden Markov model, the Viterbi algorithm is used to search for an optimal state path; The optimal state path is converted into a clean phoneme sequence based on the relationship between the state of each hidden Markov model and the basic phoneme.
7. The method of smart wake-up of a voice interactive plush toy as claimed in claim 1, wherein: The clean phoneme sequence is compared with the basic phoneme set, and based on the comparison result, the wake-up state of the plush toy is determined, and the specific steps are as follows, If the clean phoneme sequence only contains the phonemes of the wake-up word, it is determined that the plush toy enters a light wake-up state and prompts, and if the new clean phoneme sequence received within a short time window contains the phonemes of the semantic trigger word, the user is interacted with voice according to the function corresponding to the semantic trigger word, otherwise, it enters a standby state; If the clean phoneme sequence contains the phonemes of the wake-up word and the semantic trigger word, the user is interacted with voice according to the function corresponding to the semantic trigger word; If the clean phoneme sequence does not contain the phonemes of the wake-up word and the semantic trigger word, no interaction is performed, and the standby state is maintained.
8. The method of smart wake-up of a voice interactive plush toy as claimed in claim 1, wherein: The original speech signal is synchronously collected by arranging a double microphone array on the left and right sides of the plush toy. The original speech signal includes a left original speech signal collected by a left microphone and a right original speech signal collected by a right microphone. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is characterized in that: The processor executes the computer program to implement the steps of the intelligent wake-up method of the voice interaction plush toy according to any one of claims 1-8.
10. A computer readable storage medium having stored thereon a computer program, characterized in that: The computer program is executed by the processor to implement the steps of the intelligent wake-up method of the voice interaction plush toy according to any one of claims 1-8.
Citation Information
Patent Citations
Speech enhancement method and device based on dual-channel neural network time-frequency masking, and hearing-aid equipment
CN114078481A
Monaural Noise Suppression Based on Computational Auditory Scene Analysis
US20120010881A1
Methods and apparatuses for signal analysis
US6745155B1