Directional sound pickup method and device in strong interference environment, equipment and medium

By combining a super-directional beamformer, a delay-summing beamformer, and a first-level noise reduction network in a strong interference environment, the signal quality and reliability problems of traditional directional pickup technology in complex noise environments are solved, and efficient speech signal enhancement and noise suppression are achieved.

CN118887969BActive Publication Date: 2025-11-18BEIJING UNISOUND INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411052142.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2025-11-18
Estimated Expiration
2044-08-01

AI Technical Summary

Technical Problem

In environments with strong interference, traditional directional sound pickup techniques are difficult to effectively improve the quality and reliability of audio signals, especially in the presence of multipath effects and complex noise, where the accuracy of sound source localization and the performance of adaptive filtering algorithms are limited.

Method used

A super-directing beamformer and a delayed summing beamformer are combined with a pre-trained first-level noise reduction network. Through reverberation suppression, direction of arrival estimation, and adaptive canceller filters, the filter parameters are adjusted in real time to suppress interference signals from non-target directions.

Benefits of technology

It significantly improves the clarity and reliability of voice signals, and the adaptive mechanism maintains high performance stability in different environments, effectively suppressing noise and interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118887969B_ABST
    Figure CN118887969B_ABST
Patent Text Reader

Abstract

The application discloses a directional sound pickup method and device in a strong interference environment, equipment and a medium. The reverberation suppression signal is preliminarily extracted through super-directive beam forming, and then the non-human noise is removed through a first neural network noise reduction, and the speech signal mask is estimated, the speech signal mask can improve the direction of arrival estimation accuracy, so that the filter update control of the adaptive canceller is more accurate. The parameters of the adaptive canceller filter are determined according to the target sound pickup direction and the direction of arrival estimation of the speech signal, and the scheme can adjust the filter in real time to better suppress the interference signal in the non-target direction. The adaptive mechanism makes the method maintain high performance stability in different environments and conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of speech signal processing and deep learning technology, and in particular to a method, apparatus, device and medium for directional sound pickup in a strong interference environment. Background Technology

[0002] In complex acoustic environments, especially in the presence of strong interference sources (such as machine noise, crowd noise, etc.), traditional omnidirectional sound pickup methods often struggle to effectively extract clear audio signals from the target sound source. This limitation significantly impacts the effectiveness of applications in fields such as voice communication, speech recognition, and audio recording.

[0003] Directional sound pickup technology, as an effective solution, utilizes microphone array beamforming technology and adaptive filtering algorithms to achieve high-precision localization and effective pickup of target sound sources in complex acoustic environments. This technology enhances the reception of the target signal by adjusting the directivity of the pickup system to align the pickup direction with the target sound source, while simultaneously suppressing noise and interference from other directions.

[0004] However, traditional directional sound pickup techniques still face many challenges in environments with strong interference. For example, the accuracy of sound source localization may be affected by factors such as multipath effects and environmental reflections; the performance of adaptive filtering algorithms may also be limited by factors such as signal non-stationarity and complex and variable noise characteristics. Therefore, it is necessary to develop a more advanced and robust directional sound pickup method and device to cope with the complex acoustic challenges in environments with strong interference and improve the quality and reliability of audio signal pickup. Summary of the Invention

[0005] This application provides a method, apparatus, device, and medium for directional sound pickup in strong interference environments, which solves the problem of poor quality and reliability of audio signal pickup in existing speech pickup methods under strong interference environments.

[0006] In a first aspect, this application provides a method for directional sound pickup in a strong interference environment, the method comprising:

[0007] The microphone array received signal is subjected to reverberation suppression processing to obtain a reverberation-suppressed signal;

[0008] Based on the topology of the microphone array and the target pickup direction, determine the super-directional beamformer and the delay summation beamformer;

[0009] The reverberation suppression signal is processed by the super-directive beamformer to obtain the super-directive enhancement signal;

[0010] The super-directional enhancement signal is denoised using a pre-trained first-level denoising network to obtain the speech signal mask corresponding to the first-level enhancement signal and the reverberation suppression signal.

[0011] Based on the speech signal mask and the reverberation suppression signal, the direction of arrival (DOA) is estimated to obtain the speech signal DOA estimate θ.

[0012] The reverberation suppression signal is weighted by the delay summing beamformer, and the far-end signal is determined based on the weighting result.

[0013] Based on the target pickup direction and the direction of arrival of the speech signal, θ is estimated, and the filter parameters in the adaptive canceller filter are determined.

[0014] The adaptive canceller output signal is determined based on the far-end signal and the first-stage enhancement signal using the adaptive canceller filter.

[0015] The final output signal is determined based on the output signal of the adaptive canceller.

[0016] Secondly, this application also provides a directional sound pickup device for use in environments with strong interference, the device comprising:

[0017] The reverberation suppression unit is used to perform reverberation suppression processing on the signal received by the microphone array to obtain a reverberation suppressed signal;

[0018] The determining unit is used to determine the super-directional beamformer and the delay summation beamformer based on the topology of the microphone array and the target pickup direction.

[0019] A super-directional beamformer processing unit is used to process the reverberation suppression signal through the super-directional beamformer to obtain a super-directional enhancement signal;

[0020] A primary noise reduction unit is used to perform noise reduction processing on the super-directional enhancement signal through a pre-trained primary noise reduction network to obtain the speech signal mask corresponding to the primary enhancement signal and the reverberation suppression signal.

[0021] The direction of arrival estimation unit is used to perform direction of arrival estimation based on the speech signal mask and the reverberation suppression signal to obtain the speech signal direction of arrival estimation θ.

[0022] The delay summation beamformer processing unit is used to perform weighted processing on the reverberation suppression signal through the delay summation beamformer, and determine the far-end signal based on the weighting result;

[0023] The parameter determination unit is used to determine the filter parameters in the adaptive canceller filter based on the target pickup direction and the direction of arrival of the speech signal estimated θ.

[0024] An adaptive canceller filter processing unit is used to determine the adaptive canceller output signal based on the far-end signal and the first-stage enhancement signal through the adaptive canceller filter.

[0025] The integrated processing unit is used to determine the final output signal based on the output signal of the adaptive canceller.

[0026] Thirdly, this application provides a computer device including a processor, which executes a computer program stored in a memory to implement the steps of the directional sound pickup method in a strong interference environment as described above.

[0027] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the directional sound pickup method in a strong interference environment as described above.

[0028] The beneficial effects of this application are as follows:

[0029] 1. The reverberation suppression signal is initially extracted by super-directive beamforming, and then non-human noise is removed by a first-level neural network noise reduction. At the same time, the speech signal mask is estimated. The speech signal mask can improve the accuracy of the direction of arrival estimation, thereby making the filter update control of the adaptive canceller more accurate.

[0030] 2. By determining the parameters of the adaptive canceller filter based on the target pickup direction and the direction of arrival (DOA) of the speech signal, this scheme can adjust the filter in real time to better suppress interference signals from non-target directions. This adaptive mechanism enables the method to maintain high performance stability under different environments and conditions. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This application provides a schematic diagram of a directional sound pickup process in a strong interference environment.

[0033] Figure 2 This application provides an illustration of an application scenario.

[0034] Figure 3 This application provides a specific schematic diagram of a directional sound pickup process in a strong interference environment for the embodiments of this application;

[0035] Figure 4 This application provides a schematic diagram of the structure of a directional sound pickup device in a strong interference environment.

[0036] Figure 5 This is a schematic diagram of the structure of a computer device provided as an optional embodiment of this application. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0038] To address the complex acoustic challenges in environments with strong interference and improve the quality and reliability of audio signal pickup, this application provides a method, apparatus, device, and medium for directional sound pickup in environments with strong interference.

[0039] Example 1:

[0040] This application provides a method for directional sound pickup in a strong interference environment. Figure 1 This application provides a schematic diagram of a directional sound pickup process in a strong interference environment, which includes:

[0041] S101: Perform reverberation suppression processing on the microphone array received signal to obtain a reverberation suppressed signal.

[0042] In this application, the directional sound pickup method in a strong interference environment is applied to a computer device, which can be a smart terminal, such as a computer or a robot, or a server, such as an application server or a business server.

[0043] Figure 2This application provides an illustration of an application scenario. A microphone array is a system composed of multiple microphones arranged according to certain rules, capable of capturing sound signals from different directions. This technology is widely used in speech recognition, sound source localization, and audio enhancement. In complex environments, such as conference rooms and concert halls, sound signals are often affected by reverberation (also known as echo or multipath effect). This means that the sound signal is reflected multiple times by obstacles such as walls and ceilings before reaching the microphone, resulting in the received signal containing the superposition of the original sound and multiple delayed and attenuated reflected sounds. This reverberation severely reduces the clarity and recognizability of the sound signal. Based on this, in this application, after the microphone array acquires the received signal, the computer device used to execute the directional sound pickup method in a strong interference environment can perform reverberation suppression processing on the received signal of the microphone array to obtain a reverberation-suppressed signal. The computer device can employ techniques including, but not limited to, adaptive filtering, beamforming, blind source separation, and deep learning to achieve reverberation suppression processing on the received signal of the microphone array.

[0044] In one possible implementation, the reverberation suppression processing of the microphone array received signal to obtain a reverberation-suppressed signal includes:

[0045] Acquire the microphone array received signal and echo reference signal;

[0046] A short-time Fourier transform is performed on the microphone array received signal and the echo reference signal to obtain a short-time time-frequency domain signal;

[0047] The short-time frequency domain signal is subjected to reverberation suppression processing to obtain the reverberation suppressed signal.

[0048] In this application, the computer device can simultaneously acquire multiple audio signals captured by the microphone array (i.e., the microphone array received signals) and a possible echo reference signal. The echo reference signal can be a known noise source signal (such as audio played by a speaker) or a signal estimated in some way that reflects the environmental echo characteristics. In some cases, if the environmental echo characteristics are difficult to obtain directly, an algorithm can be used to generate an approximate echo model as a reference. To process the signal in the time-frequency domain, a Short-Time Fourier Transform (STFT) is required on both the microphone array received signal and the echo reference signal. STFT is a method for converting a signal from the time domain to the time-frequency domain. It divides the signal into multiple short-time windows and performs a Fourier transform on the signal within each window. Thus, the signal at each time point can be represented as a set of frequency components, facilitating subsequent processing. After obtaining the short-time time-frequency domain signal, the next step is to perform reverberation suppression processing on the short-time time-frequency domain signal. This step can utilize a variety of techniques, including but not limited to adaptive filtering, beamforming, blind source separation, and deep learning techniques. The specific choice depends on the actual application scenario and performance requirements, and is not specifically limited here.

[0049] S102: Determine the super-directional beamformer and the delay summation beamformer based on the topology of the microphone array and the target pickup direction.

[0050] In microphone array signal processing, beamforming is a key technology used to enhance sound signals from a specific direction (i.e., the direction of the target sound source) while suppressing signals from other directions (such as noise, interference, or reverberation). Based on this, in this application, a super-directional beamformer and a delay-sum beamformer can be determined according to the topology of the microphone array (i.e., the arrangement of the microphones in space) and the current pickup direction (i.e., the direction of the sound source the system wants to enhance). These two beamformers enable precise directional sound pickup in environments with strong interference. Specifically, the super-directional beamformer optimizes the array's beam pattern to achieve extremely high directional gain in a specific direction, thereby enhancing the sound signal in that direction and suppressing noise and interference from other directions. The delay-sum beamformer compensates for the time delay of signals received by different microphones and then superimposes these signals to form a directional beam.

[0051] It should be noted that the specific methods for determining the super-directional beamformer and the delay-sum beamformer based on the microphone array topology and the target pickup direction are existing technologies and will not be elaborated here.

[0052] S103: The reverberation suppression signal is processed by the super-directive beamformer to obtain the super-directive enhancement signal.

[0053] After obtaining the super-directional beamformer and the reverberation suppression signal based on the above embodiments, the reverberation suppression signal can be input into the super-directional beamformer. The super-directional beamformer processes the reverberation suppression signal to obtain a super-directional enhancement signal, thereby enhancing the sound signal from the direction of the sound source while more effectively suppressing interference signals from other directions.

[0054] S104: The super-directional enhancement signal is denoised using a pre-trained first-level denoising network to obtain the speech signal mask corresponding to the first-level enhancement signal and the reverberation suppression signal.

[0055] In the above process, after processing the reverberation suppression signal using a super-directional beamformer, a super-directional enhancement signal can be obtained. This step significantly enhances the sound signal from the direction of the target sound source and effectively suppresses interference signals from other directions, including residual reverberation and noise. Next, to further improve signal quality, noise reduction techniques can be introduced. For example, a pre-trained noise reduction network (referred to as the first-level noise reduction network) is used to denoise the super-directional enhancement signal. This noise reduction network can be built based on deep learning techniques, such as convolutional neural networks (CNN), long short-term memory networks (LSTM), or combinations thereof. Through a large amount of training data, this first-level noise reduction network can learn how to extract clean speech signals from noisy signals.

[0056] During the noise reduction process, the first-stage noise reduction network analyzes the spectral characteristics of the super-directional augmentation signal, identifying and removing noise components. Simultaneously, this first-stage network generates a speech mask, a sequence corresponding to each time frame of the signal, indicating the presence and relative intensity of the speech signal within each frame. This speech mask allows for more precise location and processing of speech signal components, further improving the noise reduction effect.

[0057] After processing by the first-stage denoising network described above, two outputs are obtained: the first-stage enhanced signal and the speech signal mask. The first-stage enhanced signal is the result of denoising the original super-directional enhanced signal, with a significantly reduced noise level and further improved speech quality. The speech signal mask provides useful information for subsequent possible signal processing steps (such as further denoising, speech coding, etc.).

[0058] S105: Based on the speech signal mask and the reverberation suppression signal, the direction of arrival (DOA) is estimated to obtain the speech signal DOA estimate θ.

[0059] In audio processing, Direction of Arrival (DOA) estimation is a crucial step used to determine the direction in which a sound signal arrives at a microphone array. Based on the speech signal mask and reverberation suppression signal obtained in the above embodiments, combined with the DOA estimation algorithm, more accurate DOA estimation can be achieved. The speech signal mask provides important clues about which time frames contain the speech signal. Since the reverberation suppression signal has removed most of the environmental reverberation and noise, combining it with the speech signal mask allows for more accurate focusing on the time periods and frequency components containing clean speech signals.

[0060] S106: The reverberation suppression signal is weighted by the delay summing beamformer, and the far-end signal is determined based on the weighting result.

[0061] In audio processing, delay-sum beamformers are used to enhance signals from a specific direction by adjusting the relative delay and weight of the signals received by each microphone in a microphone array. After obtaining the reverberation suppression signal and the delay-sum beamformer based on the above embodiment, the reverberation suppression signal can be input into the delay-sum beamformer. The delay-sum beamformer processes this reverberation suppression signal and uses it to determine the far-end signal (i.e., the signal from a specific direction, which may indicate that the actual location of the sound source is far away).

[0062] For example, the delay summation beamformer weights in the target pickup direction are weighted and applied to the reverberation suppression signals corresponding to multiple microphones respectively. These weighted reverberation suppression signals are then subtracted pairwise from the reverberation suppression signal corresponding to the reference microphone to obtain the far-end signal. The reverberation suppression signal corresponding to the reference microphone is the weighted reverberation suppression signal of the reference microphone among the multiple microphones.

[0063] S107: Based on the target pickup direction and the direction of arrival of the speech signal, estimate θ and determine the filter parameters in the adaptive canceller filter.

[0064] After obtaining the target pickup direction and the estimated direction of arrival (DOA) θ of the speech signal based on the above embodiments, the filter parameters in the Adaptive Noise Canceller can be optimized and adjusted according to the target pickup direction and the known DOA estimation θ. This adaptive canceller is used to remove unwanted noise (such as background noise, echo, etc.) from the received signal to further improve the signal-to-noise ratio. For example, based on the obtained DOA estimation θ of the speech signal and the information of the target pickup direction, the approximate location of the noise source relative to the microphone array can be inferred. Then, based on this direction and location information, the adaptive canceller adjusts the parameters of its internal filters. These parameters include the filter coefficients, order, etc., which together determine the filter's response characteristics to the input signal. Through continuous iteration and adjustment, the adaptive canceller can gradually learn and adapt to the characteristics of the noise, thereby effectively removing noise from the received signal while preserving as much useful speech signal as possible.

[0065] In one possible implementation, determining the filter parameters in the adaptive canceller filter based on the target pickup direction and the estimated direction of arrival (DOA) of the speech signal θ includes:

[0066] If the absolute value of the difference between the estimated direction of arrival (DOA) θ of the speech signal and the target pickup direction is less than a preset range threshold, then the parameters of the adaptive cancellation filter will not be updated.

[0067] If the absolute value of the difference between the estimated direction of arrival (DOA) θ of the speech signal and the target pickup direction is not less than a preset range threshold, then the parameters of the adaptive cancellation filter are updated.

[0068] In this application, when the absolute value of the difference between the estimated direction of arrival (DOA) θ of the speech signal and the target pickup direction is less than a preset threshold, it means that the sound source location is very close to or almost identical to the pickup direction. In this case, the speech component dominates the received signal, while the influence of unwanted noise such as background noise or echo is relatively small. Therefore, to avoid unnecessary calculations and adjustments, as well as possible introduced errors, it is reasonable to choose not to update the parameters of the adaptive cancellation filter.

[0069] Conversely, when the absolute value of the difference between the estimated direction of arrival (DOA) θ of the speech signal and the target pickup direction is not less than a preset threshold, it indicates a significant deviation between the sound source location and the pickup direction. In this case, the received signal may contain substantial background noise, echoes, or other unwanted noise. To improve the signal-to-noise ratio and enhance speech clarity and intelligibility, the parameters of the adaptive cancellation filter need to be updated based on the current direction information. By adjusting parameters such as the filter's coefficients and order, the filter can better adapt to the characteristics of noise, thereby more effectively removing noise from the received signal.

[0070] In practical applications, the selection of the preset range threshold needs to be determined based on the specific scenario and requirements. A threshold that is too small may lead to frequent parameter updates, increasing the computational burden; while a threshold that is too large may fail to respond promptly to changes in the location of the sound source, affecting the noise reduction effect.

[0071] Characterized by the direction of target sound pickup Taking the direction of arrival (DOA) estimation of a speech signal as θ and the preset range threshold as δ as an example, if... If the current speech signal comes from the pickup area, the parameters of the adaptive cancellation filter are not updated; if If the current speech signal originates from outside the pickup area, update the parameters of the adaptive cancellation filter.

[0072] S108: The adaptive canceller output signal is determined based on the far-end signal and the first-level enhancement signal through the adaptive canceller filter.

[0073] After optimizing and adjusting the filter parameters of the adaptive canceller filter based on the above embodiments, the input signal can be processed by the adaptive canceller filter to further improve the signal-to-noise ratio. The input signal mainly includes a first-stage enhancement signal and a far-end signal obtained through a delay-summing beamformer.

[0074] For example, the adaptive canceller filter analyzes and processes the far-end signal and the first-stage enhancement signal according to its current parameter settings to identify and remove unwanted noise components such as background noise and echoes, while preserving as much useful speech signal as possible. The signal processed by the adaptive canceller filter is output as the adaptive canceller output signal. This output signal significantly improves the signal-to-noise ratio, making the speech signal clearer and more intelligible, providing a better foundation for subsequent speech processing or communication.

[0075] S109: Determine the final output signal based on the output signal of the adaptive canceller.

[0076] After processing the input signal (including the far-end signal and the first-stage enhancement signal) through an adaptive canceller filter to obtain an effectively denoised adaptive canceller output signal, the final output signal can be determined based on this adaptive canceller output signal. For example, the final output signal can be obtained by performing a short-time inverse Fourier transform on the adaptive canceller output signal.

[0077] In one possible implementation, determining the final output signal based on the adaptive canceller output signal includes:

[0078] By using a pre-configured post-processing module, residual interference signals are suppressed on the adaptive canceller output signal based on the signal energies corresponding to the first-level enhanced signal and the adaptive canceller output signal, so as to obtain an optimized adaptive canceller output signal.

[0079] The final output signal is determined based on the optimized adaptive canceller output signal.

[0080] To further suppress residual interference signals and optimize the quality of the output signal, this application also includes a post-processing module to suppress residual interference signals in the adaptive canceller output signal. Therefore, after obtaining the adaptive canceller output signal based on the above embodiments, the post-processing module can further process the adaptive canceller output signal. For example, the post-processing module can obtain the signal energy of the first-level enhancement signal and the adaptive canceller output signal, namely, the first signal energy corresponding to the first-level enhancement signal and the second signal energy corresponding to the adaptive canceller output signal. By analyzing the first and second signal energies, the post-processing module can identify which parts contain more useful information (such as speech) and which parts may contain more interference (such as noise). Then, based on the first and second signal energies, the residual interference signals in the adaptive canceller output signal are suppressed to obtain an optimized adaptive canceller output signal. Subsequently, the final output signal can be determined based on the optimized adaptive canceller output signal. For example, a short-time inverse Fourier transform can be performed on the optimized adaptive canceller output signal to obtain the final output signal.

[0081] In one possible implementation, residual interference signals in the adaptive canceller output signal can be suppressed using an gain method. For example, the step of suppressing residual interference signals in the adaptive canceller output signal using a pre-configured post-processing module, based on the signal energies corresponding to the first-level enhancement signal and the adaptive canceller output signal respectively, to obtain an optimized adaptive canceller output signal, includes:

[0082] Obtain the first signal energy corresponding to the first-level enhanced signal and the second signal energy corresponding to the output signal of the adaptive canceller;

[0083] Determine the ratio of the first signal energy to the second signal energy;

[0084] Take the logarithm to the base 10 of the ratio;

[0085] The gain value is determined based on the relationship between the logarithm and a pre-configured adjustable suppression threshold;

[0086] Based on the gain value and the output signal of the adaptive canceller, the optimized output signal of the adaptive canceller is determined.

[0087] In this application, after the post-processing module obtains the energy of the first signal and the energy of the second signal, it can determine the ratio between the energy of the first signal (the energy of the reference signal) and the energy of the second signal (the energy of the adaptive canceller output signal). This ratio reflects the relative strength of the two signals in terms of energy, helping the post-processing module understand the enhancement or suppression effect of the adaptive canceller on the original signal. The energy ratio calculated above is logarithmically divided to base 10. Based on the relationship between the logarithmic ratio and a pre-configured adjustable suppression threshold, a gain value is determined. This threshold is set according to the actual application scenario and the desired output signal quality. If the ratio exceeds the threshold, it indicates that the residual interference signal in the adaptive canceller output signal is relatively small, and a smaller gain or even no gain may be needed; if the ratio is below the threshold, it indicates that the residual interference signal is large, and a larger gain is needed to further suppress the interference. Finally, the amplitude of the adaptive canceller output signal is adjusted according to the determined gain value to obtain the optimized adaptive canceller output signal.

[0088] In one possible implementation, determining the gain value based on the relationship between the logarithm and a pre-configured adjustable suppression threshold includes:

[0089] If the logarithm is greater than or equal to the adjustable suppression threshold, then the preset first value is determined as the gain value; wherein, the first value is less than 1;

[0090] If the logarithm is less than the adjustable suppression threshold, then the preset second value is determined as the gain value; wherein the second value is greater than or equal to 1.

[0091] In this application, if the logarithm is greater than or equal to the adjustable suppression threshold, it indicates that the energy of the adaptive canceller output signal is relatively high and the residual interference signal is relatively low. In this case, the post-processing module will determine a preset first value as the gain value. This first value is less than 1, meaning that the adaptive canceller output signal will be attenuated to a certain extent in subsequent processing to avoid distortion or discomfort caused by excessively strong signals. If the logarithm is less than the adjustable suppression threshold, it indicates that the energy of the adaptive canceller output signal is relatively low and the residual interference signal may be relatively high. In this case, the post-processing module will determine a preset second value as the gain value. This second value is greater than or equal to 1, used to amplify the amplitude of the adaptive canceller output signal in subsequent processing to further suppress residual interference signals and improve the quality of the output signal.

[0092] For example, the post-processing module suppresses residual interference signals in the output signal of the adaptive canceller by using a gain method, which can be expressed by the following formula:

[0093] S post =Gain*S AF ;

[0094]

[0095] Among them, S AF E represents the output signal of the adaptive canceller. NNI E represents the energy of the first signal corresponding to the first-level enhanced signal. AF The second signal energy corresponding to the output signal of the adaptive canceller is represented by Gain, which represents the gain value. post The optimized adaptive canceller output signal is represented by R, which represents the logarithm. th The inhibition threshold is adjustable, with the first value being 0.1 and the second value being 1.0.

[0096] In one example, determining the final output signal based on the optimized adaptive canceller output signal includes:

[0097] The optimized adaptive canceller output signal is denoised using a pre-configured two-stage noise reduction network to obtain a two-stage enhanced signal.

[0098] The second-level enhanced signal is subjected to a short-time inverse Fourier transform to obtain the final output signal.

[0099] To further improve the quality of the output signal and reduce residual noise and interference, this application also includes a pre-configured secondary noise reduction network to denoise the optimized adaptive canceller output signal. For example, after obtaining the optimized adaptive canceller output signal based on the above embodiments, the optimized adaptive canceller output signal can be input into the secondary noise reduction network. The secondary noise reduction network then denoises the optimized adaptive canceller output signal. Through this step, the noise and interference in the optimized adaptive canceller output signal will be further suppressed, resulting in a cleaner secondary enhanced signal. Then, the final output signal is obtained by performing a short-time Fourier inverse transform on the secondary enhanced signal.

[0100] The beneficial effects of this application are as follows:

[0101] 1. The reverberation suppression signal is initially extracted by super-directive beamforming, and then non-human noise is removed by a first-level neural network noise reduction. At the same time, the speech signal mask is estimated. The speech signal mask can improve the accuracy of the direction of arrival estimation, thereby making the filter update control of the adaptive canceller more accurate.

[0102] 2. By determining the parameters of the adaptive canceller filter based on the target pickup direction and the direction of arrival (DOA) of the speech signal, this scheme can adjust the filter in real time to better suppress interference signals from non-target directions. This adaptive mechanism enables the method to maintain high performance stability under different environments and conditions.

[0103] Example 2:

[0104] Figure 3 This application provides a specific schematic diagram of a directional sound pickup process in a strong interference environment, which includes:

[0105] S301: Acquire the signal received by the microphone array.

[0106] S302: Perform a short-time Fourier transform on the microphone received signal and the echo reference signal to obtain a short-time time-frequency domain signal.

[0107] S303: Perform reverberation suppression processing on the short-time frequency domain signal to obtain the reverberation suppressed signal.

[0108] S304: The reverberation suppression signal is processed by a super-directive beamformer to obtain a super-directive enhancement signal.

[0109] The super-directional beamformer is determined based on the topology of the microphone array and the target pickup direction.

[0110] S305: Through a pre-trained first-level noise reduction network, the super-directional enhancement signal is denoised to obtain the speech signal mask corresponding to the first-level enhancement signal and the reverberation suppression signal.

[0111] S306: Based on the speech signal mask and reverberation suppression signal, the direction of arrival (DOA) is estimated to obtain the speech signal DOA estimate θ.

[0112] S307: The reverberation suppression signal is weighted by a delay summation beamformer, and the far-end signal is determined based on the weighting result.

[0113] The delay summation beamformer is determined based on the topology of the microphone array and the target pickup direction.

[0114] S308: Estimate θ based on the target pickup direction and the direction of arrival of the speech signal, and determine the filter parameters in the adaptive canceller filter.

[0115] In one possible implementation, the filter parameters in the adaptive canceller filter are determined based on the target pickup direction and the speech signal direction of arrival (DOA) estimate θ, including:

[0116] If the absolute value of the difference between the estimated direction of arrival (DOA) θ of the speech signal and the target pickup direction is less than a preset range threshold, the parameters of the adaptive cancellation filter will not be updated.

[0117] If the absolute value of the difference between the estimated direction of arrival (DOA) θ of the speech signal and the target pickup direction is not less than a preset range threshold, then the parameters of the adaptive cancellation filter are updated.

[0118] S309: The adaptive canceller output signal is determined based on the far-end signal and the first-level enhancement signal through the adaptive canceller filter.

[0119] S310: Through a pre-configured post-processing module, based on the signal energy corresponding to the first-level enhanced signal and the adaptive canceller output signal, residual interference signal suppression is performed on the adaptive canceller output signal to obtain an optimized adaptive canceller output signal.

[0120] In one possible implementation, a pre-configured post-processing module suppresses residual interference signals in the adaptive canceller output signal based on the signal energies corresponding to the first-level enhanced signal and the adaptive canceller output signal, respectively, to obtain an optimized adaptive canceller output signal, including:

[0121] Obtain the first signal energy corresponding to the first-level enhanced signal and the second signal energy corresponding to the output signal of the adaptive canceller;

[0122] Determine the ratio of the energy of the first signal to the energy of the second signal;

[0123] The comparison value is taken as the logarithm to the base 10;

[0124] The gain value is determined based on the relationship between the logarithm and a pre-configured adjustable suppression threshold;

[0125] Based on the gain value and the adaptive canceller output signal, the optimized adaptive canceller output signal is determined.

[0126] In one possible implementation, the gain value is determined based on the relationship between the logarithm and a pre-configured adjustable suppression threshold, including:

[0127] If the logarithm is greater than or equal to the adjustable suppression threshold, then the preset first value is determined as the gain value; wherein, the first value is less than 1;

[0128] If the logarithm is less than the adjustable suppression threshold, then the preset second value is determined as the gain value; wherein the second value is greater than or equal to 1.

[0129] S311: The optimized adaptive canceller output signal is denoised using a pre-configured two-stage noise reduction network to obtain a two-stage enhanced signal.

[0130] S312: Perform short-time inverse Fourier transform on the secondary enhancement signal.

[0131] S313: Obtain the final output signal.

[0132] Example 3:

[0133] Based on the same inventive concept, this application also provides a directional sound pickup device for use in environments with strong interference. Figure 4 This application provides a schematic diagram of a directional sound pickup device for use in environments with strong interference. The device includes:

[0134] The reverberation suppression unit 41 is used to perform reverberation suppression processing on the microphone array received signal to obtain a reverberation suppression signal;

[0135] The determining unit 42 is used to determine the super-directional beamformer and the delay summation beamformer according to the topology of the microphone array and the target pickup direction;

[0136] The super-directive beamforming processing unit 43 is used to process the reverberation suppression signal through the super-directive beamforming to obtain the super-directive enhancement signal;

[0137] The first-level noise reduction unit 44 is used to perform noise reduction processing on the super-directional enhancement signal through a pre-trained first-level noise reduction network to obtain the speech signal mask corresponding to the first-level enhancement signal and the reverberation suppression signal.

[0138] The direction of arrival estimation unit 45 is used to perform direction of arrival estimation based on the speech signal mask and the reverberation suppression signal to obtain the speech signal direction of arrival estimation θ.

[0139] The delay summation beamformer processing unit 46 is used to perform weighted processing on the reverberation suppression signal through the delay summation beamformer, and determine the far-end signal based on the weighting result.

[0140] The parameter determination unit 47 is used to determine the filter parameters in the adaptive canceller filter based on the target pickup direction and the speech signal direction of arrival estimate θ.

[0141] The adaptive canceller filter processing unit 48 is used to determine the adaptive canceller output signal based on the far-end signal and the first-stage enhancement signal through the adaptive canceller filter.

[0142] The integrated processing unit 49 is used to determine the final output signal based on the output signal of the adaptive canceller.

[0143] In this embodiment, the directional sound pickup device in a strong interference environment is presented in the form of a functional module. Here, a module refers to an application-specific integrated circuit (ASIC), a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0144] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0145] Example 4:

[0146] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of this application, such as... Figure 5 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 5 Take a processor 10 as an example.

[0147] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0148] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0149] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device as shown by a landing page for an app. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0150] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0151] The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.

[0152] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the computer device, such as a touchscreen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touchscreen.

[0153] Example 5:

[0154] Based on the above embodiments, this application also provides a computer-readable storage medium storing a computer program executable by a processor. When the program runs on the processor, it causes the processor to perform the following steps:

[0155] The microphone array received signal is subjected to reverberation suppression processing to obtain a reverberation-suppressed signal;

[0156] Based on the topology of the microphone array and the target pickup direction, determine the super-directional beamformer and the delay summation beamformer;

[0157] The reverberation suppression signal is processed by the super-directive beamformer to obtain the super-directive enhancement signal;

[0158] The super-directional enhancement signal is denoised using a pre-trained first-level denoising network to obtain the speech signal mask corresponding to the first-level enhancement signal and the reverberation suppression signal.

[0159] Based on the speech signal mask and the reverberation suppression signal, the direction of arrival (DOA) is estimated to obtain the speech signal DOA estimate θ.

[0160] The reverberation suppression signal is weighted by the delay summing beamformer, and the far-end signal is determined based on the weighting result.

[0161] Based on the target pickup direction and the direction of arrival of the speech signal, θ is estimated, and the filter parameters in the adaptive canceller filter are determined.

[0162] The adaptive canceller output signal is determined based on the far-end signal and the first-stage enhancement signal using the adaptive canceller filter.

[0163] The final output signal is determined based on the output signal of the adaptive canceller.

[0164] Since the principle of the computer-readable storage medium in solving the problem is similar to that of the directional sound pickup method in a strong interference environment, the implementation of the computer-readable storage medium can be found in the embodiments of the method, and repeated details will not be repeated.

[0165] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for directional sound pickup in a strong interference environment, characterized in that, The method includes: The microphone array received signal is subjected to reverberation suppression processing to obtain a reverberation-suppressed signal; Based on the topology of the microphone array and the target pickup direction, determine the super-directional beamformer and the delay summation beamformer; The reverberation suppression signal is processed by the super-directive beamformer to obtain the super-directive enhancement signal; The super-directional enhancement signal is denoised using a pre-trained first-level denoising network to obtain the speech signal mask corresponding to the first-level enhancement signal and the reverberation suppression signal. Based on the speech signal mask and the reverberation suppression signal, the direction of arrival (DOA) is estimated to obtain the speech signal DOA estimate θ. The reverberation suppression signal is weighted by the delay summing beamformer, and the far-end signal is determined based on the weighting result. Based on the target pickup direction and the direction of arrival of the speech signal, θ is estimated, and the filter parameters in the adaptive canceller filter are determined. The adaptive canceller output signal is determined based on the far-end signal and the first-stage enhancement signal using the adaptive canceller filter. The final output signal is determined based on the output signal of the adaptive canceller.

2. The method as described in claim 1, characterized in that, The step of estimating θ based on the target pickup direction and the direction of arrival of the speech signal, and determining the filter parameters in the adaptive canceller filter, includes: If the absolute value of the difference between the estimated direction of arrival (DOA) θ of the speech signal and the target pickup direction is less than a preset range threshold, then the parameters of the adaptive canceller filter will not be updated. If the absolute value of the difference between the estimated direction of arrival (DOA) θ of the speech signal and the target pickup direction is not less than a preset range threshold, then the parameters of the adaptive canceller filter are updated.

3. The method as described in claim 1, characterized in that, The step of determining the final output signal based on the output signal of the adaptive canceller includes: By using a pre-configured post-processing module, residual interference signals are suppressed on the adaptive canceller output signal based on the signal energies corresponding to the first-level enhanced signal and the adaptive canceller output signal, so as to obtain an optimized adaptive canceller output signal. The final output signal is determined based on the optimized adaptive canceller output signal.

4. The method as described in claim 3, characterized in that, The step of using a pre-configured post-processing module to suppress residual interference signals in the adaptive canceller output signal based on the signal energies corresponding to the first-level enhanced signal and the adaptive canceller output signal, to obtain an optimized adaptive canceller output signal, includes: Obtain the first signal energy corresponding to the first-level enhanced signal and the second signal energy corresponding to the output signal of the adaptive canceller; Determine the ratio of the first signal energy to the second signal energy; Take the logarithm to the base 10 of the ratio; The gain value is determined based on the relationship between the logarithm and a pre-configured adjustable suppression threshold; Based on the gain value and the output signal of the adaptive canceller, the optimized output signal of the adaptive canceller is determined.

5. The method as described in claim 4, characterized in that, Determining the gain value based on the relationship between the logarithm and a pre-configured adjustable suppression threshold includes: If the logarithm is greater than or equal to the adjustable suppression threshold, then the preset first value is determined as the gain value; wherein, the first value is less than 1; If the logarithm is less than the adjustable suppression threshold, then the preset second value is determined as the gain value; wherein the second value is greater than or equal to 1.

6. The method as described in claim 3, characterized in that, The determination of the final output signal based on the optimized adaptive canceller output signal includes: The optimized adaptive canceller output signal is denoised using a pre-configured two-stage noise reduction network to obtain a two-stage enhanced signal. The second-level enhanced signal is subjected to a short-time inverse Fourier transform to obtain the final output signal.

7. The method as described in claim 1, characterized in that, The process of performing reverberation suppression processing on the microphone array received signal to obtain a reverberation suppressed signal includes: Acquire the microphone array received signal and echo reference signal; A short-time Fourier transform is performed on the microphone array received signal and the echo reference signal to obtain a short-time time-frequency domain signal; The reverberation suppression signal is obtained by performing reverberation suppression processing on the short-time frequency domain signal.

8. A directional sound pickup device for use in environments with strong interference, characterized in that, The device includes: The reverberation suppression unit is used to perform reverberation suppression processing on the signal received by the microphone array to obtain a reverberation suppressed signal; The determining unit is used to determine the super-directional beamformer and the delay summation beamformer based on the topology of the microphone array and the target pickup direction. A super-directional beamformer processing unit is used to process the reverberation suppression signal through the super-directional beamformer to obtain a super-directional enhancement signal; A primary noise reduction unit is used to perform noise reduction processing on the super-directional enhancement signal through a pre-trained primary noise reduction network to obtain the speech signal mask corresponding to the primary enhancement signal and the reverberation suppression signal. The direction of arrival estimation unit is used to perform direction of arrival estimation based on the speech signal mask and the reverberation suppression signal to obtain the speech signal direction of arrival estimation θ. The delay summation beamformer processing unit is used to perform weighted processing on the reverberation suppression signal through the delay summation beamformer, and determine the far-end signal based on the weighting result; The parameter determination unit is used to determine the filter parameters in the adaptive canceller filter based on the target pickup direction and the direction of arrival of the speech signal estimated θ. An adaptive canceller filter processing unit is used to determine the adaptive canceller output signal based on the far-end signal and the first-stage enhancement signal through the adaptive canceller filter. The integrated processing unit is used to determine the final output signal based on the output signal of the adaptive canceller.

9. A computer device, characterized in that, The computer device includes a processor that executes a computer program stored in a memory to implement the steps of the directional sound pickup method in a strong interference environment as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, It stores a computer program executable by a computer device, which, when run on the computer device, causes the computer device to perform the steps of the directional sound pickup method in a strong interference environment as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Two-dimensional directional pickup method and device

    CN113050035A

  • Beam forming method and system and beam former

    CN115086836A