Voice wake-up interaction method and system based on microphone array

By using a microphone array-based voice wake-up method, which combines VAD and sound source localization with fixed beamforming and GSC module processing, the problem of low voice wake-up rate under noise interference is solved, and efficient voice wake-up and recognition are achieved in noisy environments.

CN120808776APending Publication Date: 2025-10-17PANOVASIC TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510981516.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing voice wake-up systems have low wake-up rates and high speech recognition error rates in noisy environments, especially when the signal-to-noise ratio is extremely low or there is interference, the wake-up failure rate increases significantly.

Method used

A microphone array-based voice wake-up method is adopted. By acquiring multi-channel audio signals in real time, performing VAD processing and sound source localization, and combining pre-designed fixed beam parameters, appropriate beamforming processing is selected. The GSC module is used to enhance speech and noise, and finally voice wake-up is performed.

Benefits of technology

It effectively improves the success rate of voice wake-up, reduces the false wake-up rate, and enhances the performance of backend voice interaction tasks, especially in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808776A_ABST
    Figure CN120808776A_ABST
Patent Text Reader

Abstract

The invention discloses a voice wake-up interaction method and system based on a microphone array, and the method comprises the steps: carrying out VAD processing, so as to judge whether a target audio segment has voice or not; voice and noise source directions are obtained; selecting beam parameters of voice and noise in combination with a pre-designed fixed beam; whether GSC module processing is carried out or not is selected according to the difference between the beam parameters of the voice and the noise, so that an enhanced audio signal is obtained, a wake-up task is carried out, and a final voice wake-up result is obtained; and according to whether the wake-up is successful, determining whether to lock the voice beam direction in the current interaction stage for enhancement, thereby preventing interference of voice in other directions on subsequent interaction tasks. According to the voice enhancement mode based on the microphone array, sound source positioning can be realized, interference in a non-target direction can be suppressed, the voice quality in the target direction can be improved, the wake-up success rate can be effectively improved when the voice enhancement mode is applied to voice wake-up, and then the experience of back-end voice interaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech processing and array signal processing, in particular to a speech wake-up interaction method and system based on a microphone array. BACKGROUND

[0002] The ideal audio signal is clean and lossless, at this time, the corresponding speech wake-up, speech recognition and other interaction tasks are very easy. However, due to the complex real environment, the collected audio signal often contains noise and is unstable, resulting in low wake-up rate and high speech recognition error rate. Therefore, front-end enhancement processing is needed to improve the performance of the back-end speech corresponding task.

[0003] Most of the current speech wake-up is based on single microphone processing, when the signal-to-noise ratio is very low or there is interference, the wake-up failure rate is greatly improved. SUMMARY

[0004] The purpose of the present application is to provide a speech wake-up interaction method and system based on a microphone array, which solves the problem of low speech wake-up rate when there is noise interference in the environment.

[0005] The present application solves the above problems by the following technical solutions:

[0006] A speech wake-up interaction method based on a microphone array, comprising the following steps:

[0007] S101. Real-time acquisition of multi-channel audio, and conversion to frequency domain to obtain frequency domain signals of each channel;

[0008] S102. VAD processing according to the time domain signal or the frequency domain signal to determine whether the target audio segment contains speech;

[0009] S103. According to whether the speech exists or not, and the sound source positioning result of the frequency domain signal or the time domain signal, the speech and noise sound source direction is obtained;

[0010] S104. According to the speech and noise sound source direction, the speech and noise beam parameters are selected in combination with the pre-designed fixed beam;

[0011] S105. According to the similarity and difference of the speech and noise beam parameters, it is determined whether to perform GSC module processing; if the same, the fixed beam processing is directly used, if different, the GSC module processing is performed to obtain the enhanced audio signal;

[0012] S106. According to the enhanced audio signal after beam processing, the wake-up task is performed, and the speech wake-up is performed by comprehensively extracting the main microphone audio signal to obtain the final speech wake-up result;

[0013] S107. According to whether the wake-up is successful or not, it is determined whether to lock the current interaction stage voice beam direction for enhancement to prevent the interference of the voice in other directions on the subsequent interaction task.

[0014] As a further improvement of the present application, in S101, the specific method is:

[0015] For the real-time input of each channel time domain signal, short-time Fourier transform is performed to obtain the frequency domain signal of each channel, and the main microphone audio signal is extracted for voice wake-up.

[0016] As a further improvement of the present application, in S102, the VAD processing adopts a threshold discrimination method based on the short-time zero-crossing rate, short-time energy or spectral flatness of the time-frequency domain features, a GMM or HMM algorithm based on model matching, or a VAD method based on a neural network.

[0017] As a further improvement of the present application, in S103, the specific method is:

[0018] When the voice exists, the obtained sound source positioning result is the voice sound source direction A; when the voice does not exist, the obtained sound source positioning result is the noise sound source direction B.

[0019] As a further improvement of the present application, the method for realizing the sound source positioning adopts a cross-correlation method based on time domain signals, an adaptive filtering method, and a subspace decomposition method or a beam forming method based on frequency domain signals.

[0020] As a further improvement of the present application, the method for realizing the sound source positioning adopts GCC-PHAT and SRP-PHAT algorithms, and the specific steps include:

[0021] S301. According to the received signals of different microphones, the signals are converted to the frequency domain to obtain frequency domain signals; the specific process is as follows:

[0022] The microphone receives signals, and the signals of each channel are converted to the frequency domain to obtain frequency domain signals The specific formula is:

[0023] Formula (1);

[0024] Wherein, , represents the sound source signal of the i-th microphone at time t; t represents a certain time; is the time delay of different microphones to the reference microphone; * represents convolution; represents the transfer function between the i-th microphone receiving the sound source signal and the microphone; represents additive noise;

[0025] S302. According to the obtained frequency domain signal, the GCC-PHAT algorithm is used to estimate the time delay of the received signal between different microphones and the reference microphone in the array structure, so as to determine the range of the sound source angle; the specific formula is:

[0026] According to the frequency domain signal, the cross power spectral density is calculated:

[0027] Formula (2)

[0028] Wherein, represents the frequency domain signal of a certain frequency point; represents the conjugate of the frequency domain signal of a certain frequency point;

[0029] GCC introduces a frequency domain weighting function Enhance the robustness of time delay estimation, the weighting function adopts the PHAT method, and the specific value is:

[0030] Formula (3);

[0031] Wherein, is a very small positive value to avoid division by zero error; represents the azimuth estimated by the PHAT method;

[0032] Further processing obtains the time domain correlation function, and the maximum value thereof is obtained. Time delay and angle estimation ,

[0033] Formula (4);

[0034] Formula (5);

[0035] Wherein, represents the generalized time domain correlation function; represents the inverse Fourier transform; represents the azimuth cosine value calculated according to the above; azimuth estimated by the ij microphone; L is the array element spacing, and c is the sound propagation speed;

[0036] Finally, the final GCC-PHAT estimated angle range is obtained according to the array structure;

[0037] S303. Since the angle estimated by GCC has a certain accuracy error, the general SRP-PHAT algorithm is used to search in the target error range in the subsequent process, so as to determine the final angle.

[0038] ​As a further improvement of the present invention, in S104, the specific method is as follows: based on the VAD result and the sound source direction estimation, combined with the pre-designed fixed beams of M (M≥2), the speech beam A' and the noise beam B' are selected; when selecting the beam, it is necessary to select according to which beam range the sound source direction falls within; assuming that M beams are designed and evenly distributed from 0° to 360°, and the estimated target sound source direction H falls within the first intervals, the i-th beam is selected as the target beam.

[0039] As a further improvement of the present invention, in S105, the specific method is as follows: if the beams corresponding to the speech source direction A and the noise source direction B are the same, the fixed beam A' is directly used for processing; if the beams corresponding to the speech source direction A and the noise source direction B are different, the GSC module is processed to obtain the noisy speech and speech noise after fixed beam forming, respectively, and finally the noise is estimated by adaptively updating the ANC matrix parameters and the cancellation module is processed to obtain the final enhanced speech audio;

[0040] Among them, for the input The frame multi-channel frequency domain signal and speech beam A' are processed by FBF_1, and the noise beam B' is processed by FBF_2 to obtain the noisy speech estimation. and speech noise estimation ,after After ANC matrix filtering, the noise estimation is obtained , and finally and Perform cancellation processing to obtain the final speech estimation;

[0041] Among them, the ANC matrix weight is obtained by NLMS update, as follows:

[0042] Formula (6);

[0043] in, and Represent the current moment weight and the predicted next moment weight respectively; represents the step length, is the input signal in the description, is the output signal of the iteration;

[0044] In order to reduce speech leakage, the above weights are updated only when speech is not present, and the weight parameters remain unchanged when speech is present.

[0045] As a further improvement of the present invention, in S106, the specific method includes:

[0046] According to different environmental noise intensity scenes, different weights are given to the enhanced speech wake-up result and the main microphone speech wake-up, and finally when the wake-up value is higher than the threshold, it is considered that the wake-up is successful, and the speech wake-up result is obtained.

[0047] As a further improvement of the application, in the S107, the specific method comprises:

[0048] According to the speech wake-up success flag, it is determined whether to lock the beam direction of the current interaction stage speech, if not, the above processing continues to be performed in real time, if the wake-up is successful, the beam direction A' is taken as the speech beam direction of the current interaction stage, and the speech sound source positioning is not performed again, so as to prevent the influence of the interference speech on the target speech, until the current interaction ends and the speech sound source positioning processing is restored again.

[0049] Meanwhile, the application solves the above problems through the following technical solutions:

[0050] A speech wake-up interaction system based on a microphone array comprises:

[0051] A fixed beam design module: M fixed beam parameters are designed in advance according to the array structure and the beam number M>=2, if the array structure is regular, the corresponding parameters are directly obtained through the sound propagation principle, if the array structure is irregular, the beam parameters are designed through an offline processing mode;

[0052] A VAD module: used for determining whether the target audio segment is a speech segment, and preparing for subsequent processing;

[0053] A sound source positioning module: used for positioning the speech sound source and the noise sound source in combination with the determination result of the VAD module, and obtaining the directions of the speech and noise sound sources;

[0054] A GSC processing module: when the directions of the speech and noise sound sources are inconsistent, the GSC module is used for speech enhancement processing, mainly including a noise-containing speech first audio obtained through speech fixed beam forming processing, and a noise-containing speech second audio obtained through noise fixed beam forming processing by replacing the BM parameter estimation, and finally obtaining the final enhanced speech through adaptive updating of the ANC matrix parameters and cancellation processing;

[0055] A speech wake-up module: used for performing a speech wake-up task according to the enhanced speech, and serving as a starting end of subsequent speech interaction.

[0056] Compared with the prior art, the application has the following advantages and beneficial effects:

[0057] (1) The application utilizes a microphone array to pre-design fixed beam parameters of expected response, combines VAD processing and sound source positioning, obtains the direction of target sound source and noise source in real time, and then selects beam parameters for fixed beam forming processing, thereby reducing the amount of calculation and suppressing interference noise, so as to enhance the voice signal and improve the performance of voice wake-up under interference noise. At the same time, according to the audio positioning result when the voice wake-up is successful, the voice beam is fixed, and the performance of the interactive task such as voice recognition in the back end is improved.

[0058] (2) The voice enhancement method based on the microphone array can not only realize sound source positioning, but also suppress interference in non-target directions and improve the quality of voice in the target direction. When applied to voice wake-up, it can effectively improve the wake-up success rate and improve the experience of voice interaction in the back end. BRIEF DESCRIPTION OF DRAWINGS

[0059] Figure 1 FIG. 1 is a flowchart of a voice wake-up interactive method based on a microphone array according to the present application;

[0060] Figure 2 FIG. 2 is a sound source positioning flowchart of a voice wake-up interactive method based on a microphone array according to the present application;

[0061] Figure 3 FIG. 3 is a GSC module schematic diagram of a voice wake-up interactive method based on a microphone array according to the present application;

[0062] Figure 4 FIG. 4 is a detailed logic architecture diagram of a voice wake-up interactive method based on a microphone array according to the present application. DETAILED DESCRIPTION

[0063] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0064] Embodiment 1:

[0065] Referring to the accompanying drawings, Figures 1-4 A voice wake-up interactive method based on a microphone array mainly includes sound source positioning, beam forming and GSC algorithm processing, and specifically includes the following steps:

[0066] S101. Real-time acquisition of multi-channel audio, and conversion to the frequency domain through Fourier transform to obtain the frequency domain signals of each channel;

[0067] The specific step is: for each channel time domain signal input in real time, short time Fourier transform (STFT) is performed to obtain the frequency domain signal of each channel, and the subsequent processing is prepared. At the same time, the signal of the main microphone is directly extracted and used as a branch voice wake-up task.

[0068] S102. According to the time domain signal or the frequency domain signal, voice activity detection (VAD) processing is performed to determine whether the signal in the target length frame, that is, the target audio segment, exists voice.

[0069] The VAD processing in this step can use time domain signals or frequency domain signals; since frequency domain signals are needed later, the frequency domain is converted first, but this does not mean that the time domain signal is not used. When VAD is performed, it can also be used, which means that there are two signal methods for VAD processing.

[0070] Optionally, the VAD processing includes but is not limited to threshold discrimination methods based on time-frequency domain features such as short-time zero-crossing rate, short-time energy, or spectral flatness, and GMM (Gaussian Mixture Model) or HMM (Hidden Markov Model) algorithms based on model matching, and VAD methods based on neural networks such as silero_vad.

[0071] S103. According to whether the voice exists and the sound source positioning result of the frequency domain signal or the time domain signal, the voice sound source direction A and the noise sound source direction B are obtained.

[0072] Specifically, when the voice exists, the obtained sound source positioning result is the voice sound source direction A; when the voice does not exist, the obtained sound source positioning result is the noise sound source direction B.

[0073] Optionally, the sound source positioning method includes but is not limited to cross-correlation methods based on time domain signals such as GCC (Generalized Cross Correlation) algorithm and variants GCC-PHAT (Generalized Cross Correlation Phase Transform), adaptive filtering method, and subspace decomposition methods based on frequency domain signals such as MUSIC algorithm, beamforming methods such as MVDR, and neural network-based sound source positioning algorithms. Taking the GCC-PHAT and SRP-PHAT (Steered Response Power Phase Transform) algorithms as examples, the specific details are as follows:

[0074] S301. According to the received signals of different microphones, the signals are converted to the frequency domain to obtain frequency domain signals; the specific process is as follows:

[0075] The microphone receives a signal (the i-th microphone), and the signal of each channel is converted to the frequency domain to obtain a frequency domain signal The specific formula is:

[0076] Formula (1);

[0077] Wherein, , represents the sound source signal of the ith microphone at time t; t represents a certain time; is the time delay of different microphones to the reference microphone; * represents convolution; represents the transfer function between the sound source signal received by the ith microphone and the microphone; represents additive noise; represents the frequency domain signal of the ith microphone at time t, that is, .

[0078] S302. According to the frequency domain signal obtained in S301, the time delay of the received signal between different microphones and the reference microphone in the array structure is estimated by using the GCC-PHAT algorithm, so as to determine the range of the sound source angle. Specifically as follows:

[0079] According to the frequency domain signal, the cross power spectral density is calculated:

[0080] Formula (2)

[0081] wherein, represents the frequency domain signal of a certain frequency point; represents the conjugate of the frequency domain signal of a certain frequency point; represents the cross power spectral density calculated according to the frequency domain signal of a certain frequency point.

[0082] The GCC introduces a frequency domain weighting function to enhance the robustness of time delay estimation, and there are many kinds of weighting functions. In the present scheme, the PHAT method is adopted, and the specific value is:

[0083] Formula (3);

[0084] wherein, is a very small positive value to avoid division by zero error; represents the azimuth estimated by using the PHAT method.

[0085] Further processing obtains the time domain correlation function, and the maximum value thereof is taken to obtain the time delay and the angle estimation ,

[0086] Formula (4);

[0087] Formula (5);

[0088] wherein, represents the generalized time domain correlation function; represents the inverse Fourier transform; represents the azimuth cosine value calculated according to the foregoing; ​= represents the estimated direction based on the ij microphone; L is the array element spacing, and c is the sound propagation speed. Finally, the final GCC-PHAT estimated angle is obtained based on the array element structure. scope.

[0089] S303. Because the angle estimated by GCC has a certain precision error, the general SRP-PHAT algorithm is subsequently used to search within the target error range to determine the final angle. This will not be further described here.

[0090] S104. Based on the speech source direction A and the noise source direction B, the beam parameters of the speech and noise are selected in combination with the pre-designed fixed beam.

[0091] Specifically, based on the VAD results and the estimated sound source direction, combined with the pre-designed fixed beams of M (M ≥ 2), the speech beam A' and the noise beam B' are selected; when selecting the beam, it is necessary to select it based on which beam range the sound source direction falls within. Assume that M beams are designed, evenly distributed from 0° to 360°, and the estimated target sound source direction H falls on the first intervals, the i-th beam is selected as the target beam.

[0092] S105 selects whether to perform GSC module processing based on the similarities and differences in the beam parameters of speech and noise; if the same, then directly use the fixed beam processing, if different, then perform GSC module processing to obtain an enhanced audio signal;

[0093] If the beams corresponding to the speech source direction A and the noise source direction B are the same, beamforming is performed directly using the beam parameters corresponding to the speech source direction A, resulting in an enhanced audio output. If the beams corresponding to the speech source direction A and the noise source direction B are different, the GSC module is activated to process the enhanced speech signal. Specifically, the beam parameters A' corresponding to direction A are used to obtain the first audio signal with enhanced directionality, and the beam parameters B' corresponding to direction B are used to obtain the second audio signal with enhanced directionality. The ANC matrix parameters are then updated, and the final enhanced audio signal is obtained through the cancellation module of the GSC algorithm.

[0094] Specifically, if they are the same, the fixed beam A' is directly used for processing; if they are not the same, the GSC module is used for processing to obtain the noisy speech and speech noise after fixed beam forming, respectively. Finally, the noise is estimated by adaptively updating the ANC matrix parameters and the cancellation module is used to obtain the final enhanced speech audio.

[0095] Optionally, the GSC structure can have multiple variations. A specific architecture in the present invention is as follows: Figure 3 As shown:

[0096] For the input The frame multi-channel frequency domain signal and speech beam A' are processed by FBF_1 (fixed beam forming), and the noise beam B' is processed by FBF_2 to obtain the noisy speech estimation. and speech noise estimation ,after After ANC matrix filtering, the noise estimation is obtained , and finally and The cancellation process is performed to obtain the final speech estimation.

[0097] Among them, the ANC matrix weight is obtained by NLMS update, as follows:

[0098] Formula (6);

[0099] in, and Represent the current moment weight and the predicted next moment weight respectively; represents the step length, is the input signal in the description, is the iterative output signal; in order to reduce speech leakage, the above weights are updated only when there is no speech, and the weight parameters remain unchanged when there is speech.

[0100] S106. Perform the wake-up task based on the enhanced audio signal after beamforming, and combine it with the main microphone voice wake-up to obtain the final voice wake-up result.

[0101] Optionally, different weights can be given to the enhanced voice wake-up result and the main microphone voice wake-up result according to the different environmental noise intensity scenarios. For example,

[0102] In a noisy environment, the wake-up value = enhanced voice wake-up value 1 * 0.7 + main microphone voice wake-up value 2 * 0.3

[0103] In normal scenarios, the wake-up value = enhanced voice wake-up value 1 * 0.4 + main microphone voice wake-up value 2 * 0.6

[0104] Finally, when the wake-up value is higher than the threshold, the wake-up is considered successful, which can effectively reduce the false wake-up rate in a strong noise interference environment and improve the wake-up success rate.

[0105] S107. Based on whether the wake-up is successful or not, decide whether to lock the voice beam direction of the current interaction stage for enhancement to prevent human voices or noise from other directions from interfering with subsequent interaction tasks and improve the interaction experience.

[0106] If the wake-up is successful, the beam direction A of the current interaction stage is locked to improve the performance of subsequent interaction tasks such as speech recognition.

[0107] Optionally, according to the voice wake-up success flag, it is determined whether to lock the beam direction of the current interaction stage voice. If not woken up, the above processing continues to be performed in real time; if woken up successfully, the beam direction A is taken as the voice beam direction of the current interaction stage, and the voice sound source positioning is no longer performed, so as to prevent the influence of the interference voice on the target voice, and the noise beam is not fixed until the voice sound source positioning processing is resumed after the current interaction ends.

[0108] Optionally, if the platform with the microphone can rotate and move, the platform is rotated and moved to the front of the wake-up sound source at this time, the beam direction is locked according to the array layout, and the beam direction is taken as the target direction of the current interaction, and the noise beam direction is updated at the same time until the current interaction ends.

[0109] Embodiment 2:

[0110] A voice wake-up system based on a microphone array is used to implement the method in Embodiment 1. Through the voice enhancement method based on the microphone array, not only the sound source positioning can be realized, but also the interference in the non-target direction can be suppressed, and the voice quality in the target direction is improved. When the voice wake-up system is applied to the voice wake-up, the wake-up success rate can be effectively improved, and the experience of the back-end voice interaction is improved. The system uses the microphone array to pre-design the fixed beam parameters of the expected response, combines the VAD processing and the sound source positioning, and obtains the positions of the target sound source and the noise source in real time, and then selects the beam parameters to perform the fixed beam forming processing, so as to suppress the interference noise while reducing the calculation amount, and obtain the enhanced voice signal. At the same time, the audio signal extracted by the main microphone is also used for voice wake-up, and the final wake-up result is obtained through the two times of voice wake-up, so as to effectively improve the voice wake-up success rate and reduce the false wake-up rate. Finally, according to the audio positioning result when the voice wake-up is successful, the voice beam is fixed, and the performance of the back-end interactive tasks such as voice recognition is improved.

[0111] Specifically, a voice wake-up system based on a microphone array includes:

[0112] A fixed beam design module: M (M≥2) fixed beam parameters are pre-designed according to the array structure and the number of beams. If the array structure is regular (such as a uniform linear array, a square array, a triangular array, etc.), the corresponding parameters can be directly obtained through the sound propagation principle; if the array is irregular, the beam parameters can be designed through offline processing methods such as convex optimization algorithm and spatial response model.

[0113] A VAD module: including but not limited to the related algorithms for processing the time and frequency domain signals, and the model matching methods such as GMM, HMM, or neural network processing, to determine whether the target audio segment is a voice segment, to prepare for the subsequent processing.

[0114] Sound source positioning module: including but not limited to GCC cross-correlation algorithm, adaptive filtering algorithm, MUSIC and MVDR algorithm (any one), combined with the result of VAD, the positioning of speech sound source and noise sound source is processed, and the direction of the sound source is obtained.

[0115] GSC processing module: when the directions of speech sound source and noise sound source are inconsistent, the improved GSC module is used for speech enhancement processing. It mainly includes speech fixed beamforming processing to obtain noisy speech audio 1, and replacing BM parameter estimation with noise fixed beamforming processing to obtain noisy speech audio 2, and finally obtaining the final enhanced speech through adaptive updating of ANC matrix parameters and cancellation processing.

[0116] Voice wake-up module: according to the enhanced speech, the voice wake-up task is performed, which is the starting end of subsequent voice interaction.

[0117] Although the present application is described herein with reference to the explanatory embodiments of the present application, the above-mentioned embodiments are only the preferred embodiments of the present application, and the embodiments of the present application are not limited by the above-mentioned embodiments. It should be understood that those skilled in the art can design many other modifications and embodiments, which will fall within the scope and spirit of the principles disclosed in the present application.

Claims

1. A voice wake-up interaction method based on microphone array, characterized in that: The following steps are involved: S101 acquires multi-channel audio in real time and converts it to the frequency domain to obtain the frequency domain signal of each channel; S102 performs VAD processing based on the time domain signal or the frequency domain signal to determine whether there is speech in the target audio segment; S103. According to the presence or absence of speech, as well as the sound source localization results of the frequency domain signal or the time domain signal, obtain the direction of the speech and noise source; S104. Select the beam parameters of speech and noise according to the direction of the speech and noise sources and the pre-designed fixed beam; S105. Select whether to perform GSC module processing based on the similarities and differences in the beam parameters of speech and noise; if they are the same, then directly use fixed beam processing; if they are different, then perform GSC module processing to obtain an enhanced audio signal; S106. Perform the wake-up task based on the enhanced audio signal after beamforming, and perform voice wake-up based on the extracted main microphone audio signal to obtain the final voice wake-up result; S107. Based on whether the wake-up is successful or not, decide whether to lock the voice beam direction of the current interaction stage for enhancement to prevent human voices from other directions from interfering with subsequent interaction tasks.

2. The voice wake-up interaction method based on microphone array according to claim 1, characterized in that: In the S101, the specific method is: For the real-time input time domain signals of each channel, short-time Fourier transform is performed to obtain the frequency domain signals of each channel to prepare for subsequent processing; at the same time, the main microphone audio signal is extracted for voice wake-up.

3. The voice wake-up interaction method based on microphone array according to claim 1, characterized in that: In S102, the VAD processing adopts a threshold discrimination method based on short-time zero-crossing rate, short-time energy or spectrum flatness of time-frequency domain characteristics, a GMM or HMM algorithm based on model matching, or a VAD method based on a neural network.

4. The voice wake-up interaction method based on microphone array according to claim 1, characterized in that: In said S103, the specific method is: When speech is present, the sound source localization result obtained is the speech source direction A; when speech is not present, the sound source localization result obtained is the noise source direction B.

5. The voice wake-up interaction method based on microphone array according to claim 4, characterized in that: The method for realizing the sound source localization adopts a cross-correlation method based on time domain signals, an adaptive filtering method, and a subspace decomposition method or a beamforming method based on frequency domain signals.

6. The voice wake-up interaction method based on microphone array according to claim 4, characterized in that: The method for implementing the sound source localization adopts the GCC-PHAT and SRP-PHAT algorithms, and the specific steps include: S301. According to the received signals of different microphones, convert them into the frequency domain to obtain frequency domain signals; specifically as follows: The microphone receives the signal and converts the signal of each channel into the frequency domain to obtain the frequency domain signal The specific formula is: Formula (1); in, , represents the sound source signal of the i-th microphone at time t; t represents a certain time; is the time delay from different microphones to the reference microphone; * represents convolution; Represents the transmission function between the i-th microphone receiving the sound source signal and the microphone; represents additive noise; S302. Based on the obtained frequency domain signal, the GCC-PHAT algorithm is used to estimate the received signal delay between different microphones in the array structure and the reference microphone, thereby determining the range of the sound source angle; the specific formula is: Calculate the cross power spectral density from the frequency domain signal: Formula (2); in, Represents the frequency domain signal at a certain frequency point; Represents the conjugate of the frequency domain signal at a certain frequency point; GCC introduces frequency domain weighting function To enhance the robustness of delay estimation, the weighting function adopts the PHAT method, and its specific value is: Formula (3); in, It is a very small positive value to avoid division by zero error; Indicates the bearing estimated by the PHAT method; Further processing yields the time domain correlation function, and the maximum value is the time delay. and angle estimation , Formula (4); Formula (5); in, represents the generalized time-domain correlation function; represents the inverse Fourier transform; According to the previous calculation Azimuth cosine; represents the estimated direction based on the ij microphone; L is the array element spacing, and c is the sound propagation speed; Finally, the final GCC-PHAT estimated angle is obtained based on the array element structure scope; S303. Since the angle estimated by GCC has a certain precision error, the general SRP-PHAT algorithm is subsequently used to search within the target error range to determine the final angle.

7. The voice wake-up interaction method based on microphone array according to claim 1, characterized in that: In the above S104, the specific method is as follows: based on the VAD result and the sound source direction estimation, combined with the pre-designed fixed beams of M (M≥2), the speech beam A' and the noise beam B' are selected; when selecting the beam, it is necessary to select according to which beam range the sound source direction falls within; assuming that M beams are designed and evenly distributed from 0° to 360°, and the estimated target sound source direction H falls within the first intervals, the i-th beam is selected as the target beam.

8. The voice wake-up interaction method based on microphone array according to claim 1, characterized in that: In the S105, the specific method is as follows: if the beams corresponding to the speech source direction A and the noise source direction B are the same, then the fixed beam A' is directly used for processing; if the beams corresponding to the speech source direction A and the noise source direction B are different, then the GSC module is processed to obtain the noisy speech and speech noise after fixed beamforming, respectively, and finally the noise is estimated by adaptively updating the ANC matrix parameters and the cancellation module is processed to obtain the final enhanced speech audio; Among them, for the input The frame multi-channel frequency domain signal and speech beam A' are processed by FBF_1, and the noise beam B' is processed by FBF_2 to obtain the noisy speech estimation. and speech noise estimation ,after After ANC matrix filtering, the noise estimation is obtained , and finally and Perform cancellation processing to obtain the final speech estimation; Among them, the ANC matrix weight is obtained by NLMS update, as follows: Formula (6): in, and Represent the current moment weight and the predicted next moment weight respectively; represents the step length, is the input signal in the description, is the output signal of the iteration; In order to reduce speech leakage, the above weights are updated only when speech is not present, and the weight parameters remain unchanged when speech is present.

9. The voice wake-up interaction method based on microphone array according to claim 1, characterized in that: In S106, the specific method includes: Different weights are assigned to the enhanced voice wake-up result and the main microphone voice wake-up result based on the intensity of the ambient noise. When the wake-up value exceeds the threshold, the wake-up is considered successful and the voice wake-up result is obtained. and / or In S107, the specific method includes: Based on whether the voice wake-up success flag is set, decide whether to lock the beam direction of the voice in the current interaction phase; if not, the above processing continues in real time; if the wake-up is successful, the beam direction A' is used as the voice beam direction for this interaction phase, and voice source localization is no longer performed to prevent the interference voice from affecting the target voice. The voice source localization processing is resumed after the end of this interaction.

10. A voice wake-up interaction system based on microphone array, characterized in that: include: Fixed beam design module: pre-designs M fixed beam parameters based on the array structure and the number of beams M ≥ 2; If the array structure is regular, the corresponding parameters are directly obtained through the principle of sound propagation; if the array structure is irregular, the beam parameters are designed through offline processing; VAD module: used to determine whether the target audio segment is a speech segment and prepare for subsequent processing; Sound source localization module: used to localize the speech and noise sources based on the judgment results of the VAD module, and obtain the direction of the speech and noise sources; GSC processing module: When the directions of the speech and noise sources are inconsistent, the GSC module is used for speech enhancement. This mainly includes fixed beamforming of speech to obtain the first audio with noisy speech, replacing the BM parameter estimation with fixed beamforming of noise to obtain the second audio with speech noise, and finally, adaptively updating the ANC matrix parameters and performing cancellation processing to obtain the final enhanced speech. Voice wake-up module: used to perform voice wake-up tasks based on the enhanced voice, serving as the starting point for subsequent voice interaction.

Citation Information

Cited By

  • Underwater positioning method and positioning system based on TF-GSC acoustic signal enhancement

    CN121027997A