A doa method with a microphone array in conjunction with a wake word

By combining microphone arrays with wake words and using the DOA method, blind source separation and activation word network are used to eliminate interfering sound sources, solving the problem of inaccurate sound source localization in multi-sound-source and noisy scenes by traditional microphone arrays, and achieving more accurate sound source direction estimation.

CN117275505BActive Publication Date: 2026-08-25HANGZHOU NATCHIP SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311226945.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-20
Publication Date
2026-08-25
Estimated Expiration
2043-09-20

AI Technical Summary

Technical Problem

Traditional microphone array technology struggles to accurately estimate the direction of sound sources in multi-source and noisy environments, especially when the number of sound sources exceeds the number of array elements or when the environmental signal-to-noise ratio is low.

Method used

The DOA method, which combines a microphone array with a wake word, is used to separate the target sound source through blind source separation and activation word network, eliminate interfering human voices and noise, and estimate the sound source direction using a broadband MUSIC algorithm.

Benefits of technology

It achieves accurate estimation of the direction of the target sound source in multi-sound-source and noisy environments, avoiding the performance degradation caused by inaccurate estimation of the number of sound sources and low signal-to-noise ratio, and improving the robustness of the algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117275505B_ABST
    Figure CN117275505B_ABST
Patent Text Reader

Abstract

The application discloses a DOA method combining a microphone array with a wake-up word. First, the microphone array voice signal is separated to obtain a separated voice and noise signal; then the separated voice and noise signal is passed through an activation word network to determine the path to which the activation word belongs; the signal of the path to which the activation word belongs is excluded, and other path voices and noises are unified as interference human voice and noise reference signals; the microphone array voice signal is denoised based on the interference human voice and noise reference signals to obtain a signal sequence in which the interference human voice and noise are eliminated; a wideband MUSIC algorithm is performed on the signal sequence in which the interference human voice and noise are eliminated to obtain a final DOA, which is the human voice angle corresponding to the required activation word. The application separates multiple human voices and noises based on a blind source separation algorithm, does not need to determine the number of speakers in advance, is more robust to noise environments through noise elimination, and can obtain a more accurate DOA direction of the activation word speaker.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of speech processing technology, and in particular to microphone array sound source direction finding, specifically involving a microphone array combined with a wake word for DOA method. Background Technology

[0002] Microphone array voice direction finding (DOA) technology plays an important role in many application areas, including smart homes, audio conferencing, and personal assistants.

[0003] Traditional single-source direction finding and localization systems can only determine the location of a single sound source. With technological advancements and increasing demands, direction finding techniques for multi-source and noisy environments have emerged. Conventional multi-source speech direction finding typically involves first determining the number of sound sources, then using a microphone speech direction finding algorithm to determine the directions of multiple sources, and finally selecting the direction of the corresponding target sound source. However, conventional multi-source speech direction finding methods are affected by the number of sound sources and the environmental signal-to-noise ratio (SNR). When the number of sound sources exceeds the number of microphone array elements or the environmental SNR is low, the speech direction finding method often fails to accurately estimate the direction of each sound source, leading to algorithm failure.

[0004] Locating and tracking target sound sources in multi-source and noisy environments has always been a key focus and challenge in microphone array speech direction finding. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by providing a DOA method that combines a microphone array with a wake word. This method uses a two-step optimization strategy of width and depth to further increase the noise reduction depth while ensuring the noise reduction width.

[0006] The method of this invention is specifically as follows:

[0007] Step (1) Based on the blind source separation method, the microphone array speech signal sequence x(n)=[x1(n) … x M [s(n)] is separated to obtain the M-channel separated speech signal sequence s(n) = [s1(n) … s M [(n)], where M is the number of microphone array elements and n represents the sampling sequence number;

[0008] Step (2) takes the M-channel speech signal sequence s(n) obtained from blind source separation and passes it through an activation word network to determine the channel s to which the activation word belongs. i (n), i∈[1 M];

[0009] Step (3) is performed in s(n) = [s1(n) … s M The path s to which the speech signal containing the activation word belongs is excluded from [n]i (n), where other speech signals are unified as the reference signal sequence for interfering human voices and noise, r(n) = [s1(n) … s i-1 (n)s i+1 (n) … s M [(n)], after renumbering, is represented as r(n)=[r1(n) … r M-1 (n)];

[0010] Step (4) Denoise the microphone array speech signal sequence x(n) based on the interference voice and noise reference signal sequence r(n) to obtain a multi-channel speech signal sequence y(n) with interference voice and noise eliminated;

[0011] Step (5) uses the wideband MUSIC algorithm to obtain the final DOA of the multi-channel speech signal sequence y(n) after eliminating interference voices and noise, which is the voice angle corresponding to the required activation word.

[0012] Furthermore, the specific method of step (1) is as follows:

[0013] For the microphone array speech signal sequence x(n)=[x1(n) … x M Performing a short-time Fourier transform on [n] yields the microphone array speech signal sequence X(t,f)=[X1(t,f) … X M [(t,f)], where X m (t,f) represents the short-time Fourier transform (SFT) signal of the m=1,…,M-th microphone channel in time unit t and frequency unit f. Using the separation matrix D(t,f) of the blind source separation algorithm, M speech signal sequences S(t,f) = D(t,f)X(t,f) are obtained in the SFT domain. An inverse SFT is performed on the speech signal sequence S(t,f) in the SFT domain to obtain the M-th separated speech signal sequence s(n) = [s1(n) … s M (n)].

[0014] Furthermore, step (2) specifically involves: separating the M-channel speech signal sequence s(n) obtained from blind source separation into [s1(n) … s M [n], through the activation word network Φ, outputs the probability scoring sequence of the presence of activation words in each speech signal. Choose the one with the highest probability and score it. all the way i (n) represents the path to which the speech signal containing the activation word belongs, i∈[1 M].

[0015] Furthermore, step (4) is specifically implemented as follows: based on the reference signal sequence of interfering human voice and noise, r(n) = [r1(n)... r M-1 [x(n)], for the microphone array speech signal sequence x(n) = [x1(n) … x M Noise reduction is performed on [y1(n)] to obtain a multi-channel speech signal sequence y(n) = [y1(n) ... y2(n)] with interference from human voice and noise eliminated. M (n)]; Specifically:

[0016] First, initialize r(n) = [r1(n) ... r M-1 Each interfering human voice and noise reference signal r in [n] p (n) For the microphone array speech signal sequence x(n) = [x1(n) … x M Each microphone array speech signal x in [n] q (n) noise cancellation filter coefficients L is the order of the noise cancellation filter;

[0017] Then calculate each r p (n) for each x q The filter output of (n) The reference signal sequence of interfering human voice and noise is obtained as r(n) = [r1(n) … r M-1 [n] for the m-th microphone array speech signal x m The filter output of (n) Calculate the error signal e m (n)=x m (n)-y m (n); m = 1, ..., M;

[0018] The noise cancellation filter coefficients w are updated based on the minimum mean square error method. p,m (n+1)=w p,m (n)+μe m (n)x m (n), step size E p For r p The power of (n);

[0019] The above processing is applied to all microphone array speech signals, ultimately yielding a multi-channel speech signal sequence y(n) = [y1(n) … y2(n)] that has eliminated interfering human voices and noise. m (n) … y M (n)].

[0020] Furthermore, the specific method of step (5) is as follows: First, assume that the target direction of the activation word is θ, 0°≤θ<360°, calculate the delay of different channels corresponding to θ, and then perform delay alignment of each channel of the multi-channel speech signal sequence y(n) through a fractional delay filter to obtain the aligned multi-channel speech signal sequence y. θ (n), based on y θ (n) Calculate the covariance matrix of angle θ. The covariance matrix R(θ) is decomposed into eigenvalues ​​to obtain the corresponding eigenvectors, with the superscript H indicating the conjugate transpose; the eigenvalues ​​are sorted in order of magnitude as [b1(θ) b2(θ) … b M [θ]; Construct the cost function By iterating through different θ values, the angle corresponding to the peak value of the cost function is selected as the DOA direction of the target activation word.

[0021] The beneficial effects of this invention are: This invention separates multiple human voices and noise through a blind source separation algorithm, selects the speech path where the target sound source is located through an activation word network, and then determines other interfering sound sources and noise paths, thereby eliminating interference and noise signals in the microphone array data, and finally being able to estimate a more accurate direction of the target sound source.

[0022] The advantage of this method is that:

[0023] (1) Compared with the method of “first determine the number of sound sources and then determine the corresponding DOA of all sound sources”, it is not necessary to estimate the number of sound sources in advance, thus avoiding the decline in algorithm performance caused by inaccurate estimation of the number of sound sources.

[0024] (2) Compared with the method of “first determine the number of sound sources and then determine the corresponding DOA of all sound sources”, the traditional method requires that the number of sound sources is not greater than the number of microphone array elements. The proposed method can avoid this limitation well by using the overdetermined blind source separation algorithm. However, the algorithm is still effective when the number of sound sources is greater than the number of microphone array elements.

[0025] (3) Compared with the traditional multi-source DOA method, the traditional method is limited by the environmental signal-to-noise ratio. When the signal-to-noise ratio is too low, the algorithm performance drops sharply. However, the proposed method uses the blind source separation algorithm to separate noise and signal, which has a certain noise elimination effect. Compared with the traditional method, it is more robust to noisy scenes. Attached Figure Description

[0026] Figure 1 This is a flowchart illustrating the present invention. Detailed Implementation

[0027] To facilitate understanding of the present invention and to make the above-mentioned objects, features, and advantages of the present invention more apparent, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention, and preferred embodiments are shown in the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the disclosure of the present invention. The present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention; therefore, the present invention is not limited to the specific embodiments disclosed below.

[0028] like Figure 1 A method for DOA (Device of Orientation) combining a microphone array with a wake word is as follows:

[0029] Step (1) For the microphone array speech signal sequence x(n) = [x1(n) … x M (n)] Perform an F-point short-time Fourier transform (F = 1024 in this embodiment) to obtain the microphone array speech signal sequence X(t,f) = [X1(t,f) … X M [(t,f)], where X m (t,f) represents the short-time Fourier transform (SFT) signal of the m-th, ..., M-th microphone channel in time unit t and frequency unit f, where M is the number of microphone array elements (M=3 in this embodiment), and n represents the sampling sequence number. Using the separation matrix D(t,f) of the blind source separation algorithm, M speech signal sequences S(t,f) = D(t,f)X(t,f) are separated in the SFT domain. An inverse SFT is then performed on the speech signal sequence S(t,f) in the SFT domain to obtain the M-th separated speech signal sequence s(n) = [s1(n) ... s... M (n)].

[0030] Step (2) involves separating the M speech signal sequences s(n) obtained from blind source separation into [s1(n) … s M [n], through the activation word network Φ, outputs the probability scoring sequence of the presence of activation words in each speech signal. Choose the one with the highest probability and score it. all the way i (n) represents the path to which the speech signal containing the activation word belongs, i∈[1 M].

[0031] Step (3) is performed in s(n) = [s1(n) … s M The path s to which the speech signal containing the activation word belongs is excluded from [n] i(n), where other speech signals are unified as the reference signal sequence for interfering human voices and noise, r(n) = [s1(n) … s i-1 (n)s i+1 (n) … s M [(n)], after renumbering, is represented as r(n)=[r1(n) … r M-1 (n)].

[0032] Step (4) is based on the reference signal sequence of interfering human voice and noise, r(n) = [r1(n) … r M-1 [x(n)], for the microphone array speech signal sequence x(n) = [x1(n) … x M Noise reduction is performed on [y1(n)] to obtain a multi-channel speech signal sequence y(n) = [y1(n) ... y2(n)] with interference from human voice and noise eliminated. M (n)]; Specifically:

[0033] First, initialize r(n) = [r1(n) ... r M-1 Each interfering human voice and noise reference signal r in [n] p (n) For the microphone array speech signal sequence x(n) = [x1(n) … x M Each microphone array speech signal x in [n] q (n) noise cancellation filter coefficients L is the order of the noise cancellation filter; in this embodiment, L = 512. In this embodiment, the filter coefficients are initialized to all 1s. p,q =[1 1 …1].

[0034] Then calculate each r p (n) for each x q The filter output of (n) The reference signal sequence of interfering human voice and noise is obtained as r(n) = [r1(n) … r M-1 [n] for the m-th microphone array speech signal x m The filter output of (n) Calculate the error signal e m (n)=x m (n)-y m (n); m = 1, ..., M.

[0035] The noise cancellation filter coefficients w are updated based on the minimum mean square error method. p,m (n+1)=w p,m (n)+μe m (n)x m (n), step size E p For r pThe power of (n);

[0036] The above processing is applied to all microphone array speech signals, ultimately yielding a multi-channel speech signal sequence y(n) = [y1(n) … y2(n)] that has eliminated interfering human voices and noise. m (n) … y M (n)].

[0037] Step (5) uses the wideband MUSIC algorithm to obtain the final DOA (Directional Angle of Voice), which is the voice angle corresponding to the required activation word, from the obtained multi-channel speech signal sequence y(n) after eliminating interference voices and noise. The details are as follows:

[0038] First, assume the target direction of the activation word is θ, where 0c ≤ θ < 360°. Calculate the delay for different channels corresponding to θ, and then align the delays of each channel of the multi-channel speech signal sequence y(n) using a fractional delay filter to obtain the aligned multi-channel speech signal sequence y. θ (n), based on y θ (n) Calculate the covariance matrix of angle θ. The covariance matrix R(θ) is decomposed into eigenvalues ​​to obtain the corresponding eigenvectors, with the superscript H indicating the conjugate transpose; the eigenvalues ​​are sorted in order of magnitude as [b1(θ) b2(θ) … b M [θ]; Construct the cost function By iterating through different θ values, the angle corresponding to the peak value of the cost function is selected as the DOA direction of the target activation word.

[0039] It should be understood that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and are not intended to limit the invention. The scope of protection of this application is not limited thereto.

Claims

1. A DOA method combining a microphone array and a wake word, characterized in that: Step (1) Based on the blind source separation method, the microphone array speech signal sequence is processed. Separate to obtain Speech signal sequence after path separation ,in This refers to the number of microphone array elements. Indicates the sampling sequence number; Step (2) separates the blind source. Road speech signal sequence By using an activation word network, the path to which the speech signal containing the activation word belongs is determined. , ; Step (3) in The path to which the speech signal containing the activating word belongs is excluded. Other speech signals are used as a unified reference signal sequence for interfering human voices and noise. After being renumbered, it is represented as ; Step (4) Based on the reference signal sequence of interfering human voice and noise microphone array speech signal sequence Noise reduction was performed to obtain a multi-channel speech signal sequence with interference from human voices and noise eliminated. ; Step (5) involves processing the obtained multi-channel speech signal sequence after removing interfering human voices and noise. The final DOA is obtained through the broadband MUSIC algorithm, which is the human voice angle corresponding to the required activation word.

2. The DOA method using a microphone array combined with a wake word as described in claim 1, characterized in that: The specific method for step (1) is as follows: microphone array speech signal sequence Perform a short-time Fourier transform to obtain the microphone array speech signal sequence in the short-time Fourier transform domain. ,in Indicates time unit Frequency unit The The short-time Fourier transform domain signals of each microphone channel; the separation matrix obtained through the blind source separation algorithm. In the short-time Fourier transform domain, the separation is obtained Road speech signal sequence For speech signal sequences in the short-time Fourier transform domain Perform a short-time inverse Fourier transform to obtain the time domain. Speech signal sequence after path separation .

3. The DOA method using a microphone array combined with a wake word as described in claim 2, characterized in that: The specific method for step (2) is as follows: By activating word networks Output the probability scoring sequence of the activation words present in each speech signal. Choose the one with the highest probability and score it. all the way The path to which the speech signal containing the activation word belongs. .

4. The DOA method using a microphone array combined with a wake word as described in claim 3, characterized in that: The specific method for step (4) is as follows: Based on reference signal sequences of interfering human voice and noise For microphone array speech signal sequences Noise reduction was performed to obtain a multi-channel speech signal sequence with interference from human voices and noise eliminated. ; Specifically: First initialize Each interfering human voice and noise reference signal in microphone array speech signal sequence The voice signal of each microphone array in Noise cancellation filter coefficients , , , The order of the noise cancellation filter; Then calculate each For each Filter output The reference signal sequence of interfering human voice and noise was obtained. For the first Voice signal from a microphone array Filter output Calculate the error signal , ; Update the noise cancellation filter coefficients based on the minimum mean square error method. Step length , for The power; The above processing is performed on all microphone array speech signals to obtain a multi-channel speech signal sequence that eliminates interfering human voices and noise. .

5. The DOA method using a microphone array combined with a wake word as described in claim 4, characterized in that: The specific method for step (5) is as follows: First, let's assume the target direction of the activation word is... , ,calculate The corresponding delays for different channels, and the multi-channel speech signal sequence Each channel is time-delayed and aligned using a fractional delay filter to obtain an aligned multi-channel speech signal sequence. ,based on Calculate angle covariance matrix For the covariance matrix Perform eigenvalue decomposition to obtain the corresponding eigenvectors, with superscripts... Indicates conjugate transpose; Sort by eigenvalue size Construct the cost function traversing different The angle corresponding to the peak of the cost function is selected as the DOA direction of the target activation word. .

Citation Information

Patent Citations

  • Sound source positioning method and device, computer readable storage medium and electronic equipment

    CN112799016A

  • Voice separation method

    CN113470689A