Hearing aid omnidirectional listening speech enhancement method and system based on microphone array
By employing an omnidirectional listening speech enhancement method using microphone arrays, and utilizing the frequency domain representation of front and rear microphones and Wiener filtering technology, a combination of high speech gain and natural auditory scene is achieved in noisy environments, thereby improving the user's speech intelligibility and comfort.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUXI QINGER VOICE TECH CO LTD
- Filing Date
- 2026-02-02
- Publication Date
- 2026-05-01
AI Technical Summary
Existing beamforming technology for hearing aids struggles to achieve high speech gain in noisy environments while maintaining the naturalness of the auditory scene and the flexibility of the sound source. This results in users losing ambient sound information in noisy environments, affecting comfort and safety.
The omnidirectional listening speech enhancement method based on microphone arrays constructs an amplitude response function using the frequency domain representation of forward and backward microphones, obtains the forward narrow-beam signal and the surrounding wide-beam signal, and combines Wiener filtering technology for noise cancellation, ensuring the enhancement of the target speech signal and the suppression of environmental noise.
Without reducing speech signals from other directions, it enhances the target direction signal, improving speech intelligibility and comfort for hearing aid users, thus solving the problem of excessive suppression of ambient sounds in traditional beamforming technology.
Smart Images

Figure CN121967986A_ABST
Abstract
Description
A Method and System for Enhancing Speech in Hearing Aids Based on Microphone Arrays Technical Field
[0001] This invention relates to the field of hearing aid technology, and in particular to a method and system for enhancing omnidirectional listening speech in hearing aids based on a microphone array. Background Technology
[0002] As an important hearing aid, the core function of hearing aids is to compensate for hearing loss and improve speech clarity and intelligibility in complex everyday auditory environments. However, real-life sound scenarios often contain multiple sound sources, such as the target speech, background noise, and interfering human voices. Especially in noisy environments like restaurants and shopping malls, the signal-to-noise ratio is usually low, which poses a significant challenge to hearing aid users.
[0003] Currently, microphone array-based speech enhancement technology is widely used in modern hearing aids to improve speech quality in noisy environments. Among these, beamforming technology is the most mainstream solution. Its basic principle is to utilize the phase and amplitude differences of sound signals received by multiple microphones, and through adaptive weighting processing, form a directional pickup beam (i.e., a "sound pickup horn") in space. The main lobe of this beam points towards the target sound source, while the nulls of the beam are aligned with the direction of interfering noise, thereby enhancing the target speech and suppressing background noise. However, traditional beamforming technology has significant limitations in hearing aid applications. The highly directional beam excessively suppresses all sounds outside the target direction, resulting in an extremely unnatural auditory experience, as if listening to the world through a "sound tunnel." Users lose environmental sound information and cannot perceive the overall situation of their surroundings, which not only reduces immersion but may also pose safety hazards in certain situations (such as crossing the street).
[0004] Therefore, hearing aid users want to hear the target speaker clearly in noise, but they also do not want to lose their overall auditory perception of the surrounding environment, that is, to achieve the effect of "omnidirectional listening". However, existing technologies have difficulty ensuring high speech gain while taking into account the naturalness of the auditory scene and the flexibility of the sound source. Summary of the Invention
[0005] Therefore, it is necessary to provide a microphone array-based omnidirectional hearing aid speech enhancement method and system that can ensure high speech gain while taking into account the naturalness of the auditory scene and the flexibility of the sound source, in order to address the above-mentioned technical problems.
[0006] This invention provides a method for enhancing omnidirectional listening speech in hearing aids based on a microphone array. The method includes: determining forward and backward wide-beam signals according to the frequency domain characteristics of the forward and backward microphones, and constructing forward and backward amplitude response functions based on the forward and backward wide-beam signals; determining the time-frequency energy distribution of the microphone-acquired signal based on the ratio of the forward and backward amplitude response functions to obtain a forward narrow-beam signal and surrounding wide-beam signals; estimating the environmental noise of the forward narrow-beam signal and surrounding wide-beam signals by using the minimum power spectrum signal within a set time window to determine whether the target speech exists, and outputting the noise power spectrum; estimating the power spectrum of the target speech signal based on the noise power spectrum, and using Wiener filtering technology to perform noise cancellation to obtain the amplitude spectrum of the synthesized target speech signal after noise reduction.
[0007] In one embodiment, determining the front and rear wide-beam signals based on the frequency domain representations of the front and rear microphones, and constructing front and rear amplitude response functions based on the front and rear wide-beam signals, includes: obtaining the microphone spacing and the signal sampling rate and discrete frequency points of the microphone acquisition signals; constructing frequency domain representations of the front and rear microphones based on the microphone spacing, signal sampling rate, and discrete frequency points, combined with the speed of sound propagation in air and the angle between the target sound source and the line connecting the front and rear microphones; constructing a front amplitude response function and a rear amplitude response function based on the frequency domain representations of the front and rear microphones, and calculating the ratio between the front amplitude response function and the rear amplitude response function.
[0008] In one embodiment, determining the time-frequency energy distribution of the microphone-acquired signal based on the ratio of the forward and backward amplitude response functions to obtain a forward narrow-beam signal and a surrounding wide-beam signal includes: determining the degree of deviation of the microphone-acquired signal from the direction of the sound source in any time-frequency unit based on the magnitude of the ratio of the forward and backward amplitude response functions to determine the time-frequency energy distribution, wherein the degree of deviation of the direction of the sound source is proportional to the ratio of the forward and backward amplitude response functions; obtaining the time-frequency masking matrices of the forward narrow-beam signal and the surrounding wide-beam signal based on a set forward target angle and the ratio of the forward and backward amplitude response functions, and generating the forward narrow-beam signal and the surrounding wide-beam signal based on the time-frequency masking matrices and the frequency domain characterization of the forward microphone.
[0009] In one embodiment, estimating the environmental noise of the positive narrow-beam signal and the surrounding wide-beam signal by using the minimum power spectrum signal within a set time window to determine whether the target speech exists and outputting the noise power spectrum includes: acquiring the power spectrum signals of the positive narrow-beam signal and the surrounding wide-beam signal, and determining the minimum power spectrum signal within the set time window; calculating the ratio of the power of each channel of the current speech frame to the minimum power spectrum signal, so as to define the speech presence coefficient of the target speech in each channel of the current speech frame according to the magnitude of the ratio; wherein the speech presence coefficient is proportional to the magnitude of the ratio of the power of the corresponding channel of the current speech frame to the minimum power spectrum signal.
[0010] In one embodiment, the step of estimating the environmental noise of the positive narrow-beam signal and the surrounding wide-beam signal by using the minimum power spectrum signal within a set time window to determine whether the target speech exists and outputting the noise power spectrum further includes: binarizing the speech presence coefficient by setting a threshold, so that the presence of the target speech is determined when the speech presence coefficient exceeds the set threshold, otherwise the target speech does not exist, thus obtaining a binarized speech representation; introducing a probability smoothing factor to smooth the binarized speech representation, converting the binarized speech representation into a speech presence probability; and based on the speech presence probability, introducing a noise coefficient smoothing factor to estimate the environmental noise of the power of each channel, thereby outputting the noise power spectrum.
[0011] In one embodiment, the step of estimating the target speech signal power spectrum based on the noise power spectrum and applying Wiener filtering technology for noise cancellation to obtain the amplitude spectrum of the synthesized target speech signal after noise reduction includes: calculating the target speech signal power spectrum corresponding to the positive narrow-beam signal and the surrounding wide-beam signal based on the noise power spectra of the positive narrow-beam signal and the surrounding wide-beam signal, respectively; substituting the target speech signal power spectrum and the noise power into the Wiener filtering technology, and introducing a smoothing coefficient and a signal-to-noise ratio factor to determine the Wiener filter noise reduction coefficient.
[0012] In one embodiment, the step of estimating the power spectrum of the target speech signal based on the noise power spectrum and applying Wiener filtering technology for noise cancellation to obtain the amplitude spectrum of the synthesized target speech signal after noise reduction further includes: performing environmental noise cancellation processing on the positive narrow-beam signal and the surrounding wide-beam signal based on the Wiener filter noise reduction coefficients corresponding to the positive narrow-beam signal and the surrounding wide-beam signal respectively, to obtain the noise-cancelled positive narrow-beam signal and the surrounding wide-beam signal; and synthesizing the noise-cancelled positive narrow-beam signal and the surrounding wide-beam signal to output the amplitude spectrum of the target speech signal.
[0013] This invention also provides a microphone array-based omnidirectional hearing aid speech enhancement system, used to implement the microphone array-based omnidirectional hearing aid speech enhancement method described above. The system includes: a function construction module, used to determine forward and backward wide-beam signals respectively based on the frequency domain representation of the forward and backward microphones, and construct forward and backward amplitude response functions respectively based on the forward and backward wide-beam signals; a signal acquisition module, used to determine the time-frequency energy distribution of the microphone-acquired signal based on the ratio of the forward and backward amplitude response functions, so as to acquire a forward narrow-beam signal and surrounding wide-beam signals; a noise estimation module, used to estimate the environmental noise of the forward narrow-beam signal and surrounding wide-beam signals by using the minimum power spectrum signal within a set time window, so as to determine whether the target speech exists, and output the noise power spectrum; and a noise cancellation module, used to estimate the power spectrum of the target speech signal based on the noise power spectrum, and call Wiener filtering technology to perform noise cancellation, so as to obtain the amplitude spectrum of the synthesized target speech signal after noise reduction.
[0014] The present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the microphone array-based omnidirectional listening speech enhancement method for hearing aids as described above.
[0015] The present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, implements the microphone array-based omnidirectional listening speech enhancement method for hearing aids as described above.
[0016] The aforementioned microphone array-based omnidirectional hearing aid speech enhancement method and system determines the forward and backward wide-beam signals based on the frequency domain representations of the forward and backward microphones, and constructs forward and backward amplitude response functions based on these signals. Then, based on the ratio of the forward and backward amplitude response functions, the time-frequency energy distribution of the microphone-acquired signals is determined to obtain the forward narrow-beam signal and the surrounding wide-beam signals. Subsequently, the environmental noise of the forward narrow-beam signal and the surrounding wide-beam signals is estimated by setting the minimum power spectrum signal within a time window to determine the presence of the target speech, and the noise power spectrum is output. Finally, the power spectrum of the target speech signal is estimated based on the noise power spectrum, and Wiener filtering is used for noise cancellation to obtain the amplitude spectrum of the synthesized target speech signal after noise reduction. This invention enhances the target direction signal while suppressing environmental noise in other directions without reducing speech signals in other directions, thereby improving the maximum speech intelligibility and comfort of hearing aid users. It effectively solves the technical problem of existing beamforming technologies that only focus on the target direction signal, leading to excessive suppression of sounds outside the target direction. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 is a flowchart illustrating one of the omnidirectional listening speech enhancement methods for hearing aids based on a microphone array provided by the present invention; Figure 2 is a schematic diagram illustrating the overall process of omnidirectional listening speech enhancement in a specific embodiment of the omnidirectional listening speech enhancement method for hearing aids based on a microphone array provided by the present invention; Figure 3 is a flowchart illustrating another of the omnidirectional listening speech enhancement methods for hearing aids based on a microphone array provided by the present invention; Figure 4 is a flowchart illustrating a third of the omnidirectional listening speech enhancement methods for hearing aids based on a microphone array provided by the present invention; Figure 5 is a flowchart illustrating a fourth of the omnidirectional listening speech enhancement methods for hearing aids based on a microphone array provided by the present invention; Figure 6 is a flowchart illustrating a fifth of the omnidirectional listening speech enhancement methods for hearing aids based on a microphone array provided by the present invention; Figure 7 is a flowchart illustrating a sixth of the omnidirectional listening speech enhancement methods for hearing aids based on a microphone array provided by the present invention; Figure 8 is a flowchart illustrating a seventh of the omnidirectional listening speech enhancement methods for hearing aids based on a microphone array provided by the present invention; Figure 9 is a structural schematic diagram of the omnidirectional listening speech enhancement system for hearing aids based on a microphone array provided by the present invention; Figure 10 is an internal structural diagram of the electronic device provided by the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] The present invention relates to a microphone array-based omnidirectional listening speech enhancement method and system for hearing aids. The following description, in conjunction with Figures 1 to 10, illustrates the present invention.
[0021] As shown in Figure 1, in one embodiment, a method for enhancing omnidirectional listening speech in a hearing aid based on a microphone array includes the following steps: Step S110, determining the front and rear wide-beam signals respectively based on the frequency domain representation of the front and rear microphones, and constructing the front and rear amplitude response functions respectively based on the front and rear wide-beam signals.
[0022] Specifically, the hearing aid server determines the forward wide-beam signal and the backward wide-beam signal, i.e., the front and rear wide-beam signals, based on the frequency domain representations of the forward and rear wide-beam signals, and constructs the forward amplitude response function and the backward amplitude response function, i.e., the front and rear amplitude response functions, based on the obtained forward wide-beam signal and backward wide-beam signal, respectively.
[0023] Referring to Figure 2, in a specific embodiment, the omnidirectional listening speech enhancement method for hearing aids based on a microphone array provided by the present invention includes the following overall process: firstly, extracting forward and backward wide-beam signals from a dual-microphone array; then, obtaining a forward narrow-beam signal and surrounding wide-beam signals based on the ratio of the forward and backward wide-beam signals; and finally, denoising the forward narrow-beam signal and surrounding wide-beam signals and merging them to obtain a target speech signal with directional features.
[0024] First, acquire the forward wide-beam signal. and backward wide beam signal The expression is:
[0025]
[0026] In the formula, This is the frequency domain expression for the forward microphone; This is the frequency domain expression for the rear microphone; Indicates the spacing between the two microphones; This indicates the speed at which sound travels through the air. For discrete frequency points; The maximum value at discrete frequency points; The discrete time point number; It is half the signal sampling rate.
[0027] in, The ideal expression is:
[0028] In the formula, This indicates the angle between the target sound source and the line connecting the microphone and the two microphones.
[0029] Then, the amplitude system response function of the forward signal can be obtained separately. The system response function of the amplitude of the backward signal They are respectively:
[0030]
[0031] Step S120: Based on the ratio of the forward and backward amplitude response functions, determine the time-frequency energy distribution of the microphone-acquired signal to obtain the forward narrow-beam signal and the surrounding wide-beam signal.
[0032] Specifically, the hearing aid server determines the time-frequency energy distribution of the microphone signals acquired by the forward and backward microphones in each time-frequency unit based on the ratio between the forward amplitude response function and the backward amplitude response function obtained in step S110, so as to extract the positive narrow-beam signal and the surrounding wide-beam signal.
[0033] Referring to Figure 2, in a specific embodiment, the microphone array-based omnidirectional hearing enhancement method for hearing aids provided by this invention, during the acquisition of the forward narrow-beam signal and the surrounding wide-beam signal, has zeros of the amplitude system response functions of the forward and backward signals located between 180° and 0°, respectively. Therefore, it is first necessary to calculate the amplitude system response function of the aforementioned forward signal. The system response function of the amplitude of the backward signal The ratio between , is represented as:
[0034] Then, the energy distribution within a specific time-frequency unit of the mixed signal acquired by the microphone is determined based on the ratio of the system response to the forward and backward signal amplitudes. The closer the direction of the incoming sound source deviates from 0°, the higher the ratio. The smaller the ratio, the higher the ratio. The larger. Assume the positive target angle is... Based on the above ratio Obtain the positive narrow-beam time-frequency masking matrix respectively Time-frequency masking matrix of surrounding wide-beam signals , is represented as:
[0035]
[0036] In the formula, the parameter The range is [0,1], which represents the degree of attention to surrounding signals other than the positive target signal. The larger the value, the higher the degree of attention.
[0037] Finally, the positive narrow-beam signal can be obtained based on the aforementioned positive narrow-beam time-frequency masking matrix and the time-frequency masking matrix of the surrounding wide-beam signal. and surrounding wide beam signals , is represented as:
[0038]
[0039] Step S130: Estimate the ambient noise of the positive narrow beam signal and the surrounding wide beam signal by setting the minimum power spectrum signal within the time window to determine whether the target speech exists, and output the noise power spectrum.
[0040] Specifically, the hearing aid server estimates the ambient noise of the positive narrow beam signal and the surrounding wide beam signal by using the minimum power spectrum signal within a set time window, in order to determine whether the target speech exists in the positive narrow beam signal and the surrounding wide beam signal. Based on the determination result and the probability of the target speech, it outputs the noise power spectrum corresponding to the ambient noise.
[0041] Referring to Figure 2, in a specific embodiment, the microphone array-based omnidirectional hearing enhancement method for hearing aids provided by this invention needs to suppress environmental noise in all directions while ensuring that the quality of the speech signal is not reduced during the noise reduction processing of the forward narrow-beam signal and the surrounding wide-beam signal. First, the environmental noise of the forward and surrounding beam signals is estimated separately, assuming:
[0042] In the formula, Power spectrum signal representing any positive or surrounding beam signal.
[0043] Assuming the environmental noise is additive, meaning the power of noisy speech is the sum of the speech power and the noise power, the power of a speech frame will be significantly higher than that of a frame without speech. Furthermore, speech signals are typically not continuous; even in continuous dialogue, there will be pure noise frames without speech. Therefore, the power of a noisy frame can be determined by recording the shortest minimum power. The noise is estimated using the following formula:
[0044] In the formula, , These are all smoothing factors, used to ensure that the current frame features can be updated in real time even if the short-term minimum power depends on previous frames.
[0045] Then, with short-term minimum power As a reference for noise signals, the power of each channel in the current frame and the short-term minimum power The larger the ratio, the greater the probability that there is a speech signal in the corresponding channel of the current frame. This relationship is defined as the speech presence coefficient. Subsequently, a corresponding threshold was set for the speech presence coefficient. When the value exceeds a threshold, speech exists within the channel, and a binarized speech presence coefficient is obtained. The expression is:
[0046]
[0047] To avoid estimation noise fluctuations introduced by binarization abrupt changes, the binarized speech presence probability is smoothed to obtain the speech presence probability. The expression is:
[0048] In the formula, There is a probability smoothing factor for speech.
[0049] Finally, based on the above probability of speech existence Estimating the noise figure Furthermore, based on the noise figure Obtaining the noise power spectrum The expression is:
[0050]
[0051] Step S140: Estimate the power spectrum of the target speech signal based on the noise power spectrum, and use Wiener filtering technology to eliminate noise, thereby obtaining the amplitude spectrum of the synthesized target speech signal after noise reduction.
[0052] Specifically, the hearing aid server estimates the signal power spectrum corresponding to the target speech based on the noise power spectrum obtained in step S130, i.e., the target speech signal power spectrum. It then substitutes the noise power spectrum and the target speech signal power spectrum into the Wiener filtering technique to achieve noise cancellation of the positive narrow beam signal and the surrounding wide beam signal, and finally synthesizes and outputs the noise-reduced target speech signal amplitude spectrum.
[0053] Referring to Figure 2, in a specific embodiment, the microphone array-based omnidirectional hearing aid speech enhancement method provided by the present invention obtains the noise power spectrum. Next, the environmental estimated noise power spectra of the positive narrow-beam signal and the surrounding wide-beam signal are determined separately, denoted as and This leads to the estimated power spectrum of the positive target speech signal. Power spectrum of surrounding target speech signals The expression is:
[0054]
[0055] In the formula, It is a small value to ensure the minimum lower limit of existence for each time frequency point.
[0056] After that, set The target speech power spectrum represents the forward or surrounding beam signal, based on the estimated target speech signal power spectrum. With noise power spectrum The Wiener filtering technique is used for noise reduction. To balance noise reduction and distortion, a smoothing coefficient is introduced into the Wiener filter. and signal-to-noise ratio factor The expression is:
[0057]
[0058] In the formula, The signal-to-noise ratio (SNR) of the k-th frame is represented by the SNR factor. It changes with the signal-to-noise ratio; This represents the lower limit of the signal-to-noise ratio. This represents the upper limit of the signal-to-noise ratio; , These represent the maximum and minimum signal-to-noise ratio factors, respectively. This represents the signal-to-noise ratio factor change function when the signal-to-noise ratio value is between the upper and lower limits. This represents the noise reduction coefficient of the Wiener filter.
[0059] Finally, based on the above expressions, the Wiener filter noise reduction coefficients for the forward and surrounding beam signals can be obtained, respectively. and This allows us to obtain the final target signal amplitude spectrum synthesized from the denoised beams of the forward and surrounding beams. The expression is:
[0060] The aforementioned microphone array-based omnidirectional hearing aid speech enhancement method determines the forward and backward wide-beam signals based on the frequency domain representations of the forward and backward microphones, and constructs forward and backward amplitude response functions based on these signals. Then, based on the ratio of the forward and backward amplitude response functions, the time-frequency energy distribution of the microphone-acquired signals is determined to obtain the forward narrow-beam signal and the surrounding wide-beam signals. Subsequently, the environmental noise of the forward narrow-beam signal and the surrounding wide-beam signals is estimated by setting the minimum power spectrum signal within a time window to determine the presence of the target speech, and the noise power spectrum is output. Finally, the power spectrum of the target speech signal is estimated based on the noise power spectrum, and Wiener filtering is applied for noise cancellation to obtain the amplitude spectrum of the synthesized target speech signal after noise reduction. This method enhances the target direction signal while suppressing environmental noise in other directions without reducing speech signals in other directions, thereby improving the maximum speech intelligibility and comfort for hearing aid users. It effectively solves the technical problem of existing beamforming techniques that focus only on the target direction signal, leading to excessive suppression of sounds outside the target direction.
[0061] As shown in Figure 3, in one embodiment, the omnidirectional listening speech enhancement method for hearing aids based on microphone array provided by the present invention includes the following steps in step S110: Step S111, obtaining the microphone spacing and the signal sampling rate and discrete frequency points of the microphone acquisition signal, and constructing frequency domain representations of the front and rear microphones based on the microphone spacing, signal sampling rate and discrete frequency points, combined with the speed of sound propagation in the air and the angle between the target sound source and the line connecting the front and rear microphones.
[0062] Step S112: Based on the frequency domain representations of the forward and backward microphones, construct the forward amplitude response function and the backward amplitude response function respectively, and calculate the ratio between the forward amplitude response function and the backward amplitude response function.
[0063] As shown in Figure 4, in one embodiment, the omnidirectional listening speech enhancement method for hearing aids based on a microphone array provided by the present invention includes the following steps in step S120: Step S121, determining the degree of deviation of the sound source direction of the microphone-collected signal in any time-frequency unit according to the ratio of the forward and backward amplitude response functions, so as to determine the time-frequency energy distribution. The degree of deviation of the sound source direction is proportional to the ratio of the forward and backward amplitude response functions.
[0064] Step S122: Based on the set forward target angle and the ratio of the forward and backward amplitude response functions, the time-frequency masking matrices of the forward narrow beam signal and the surrounding wide beam signal are obtained respectively, and the forward narrow beam signal and the surrounding wide beam signal are generated according to the time-frequency masking matrices and the frequency domain representation of the forward microphone.
[0065] As shown in Figure 5, in one embodiment, the omnidirectional listening speech enhancement method for hearing aids based on a microphone array provided by the present invention includes the following steps in step S130: Step S131, acquiring the power spectrum signals of the positive narrow beam signal and the surrounding wide beam signal, and determining the minimum power spectrum signal within a set time window.
[0066] Step S132: Calculate the ratio of the power of each channel in the current speech frame to the minimum power spectrum signal, so as to define the speech presence coefficient of each channel in the current speech frame that contains the target speech based on the ratio.
[0067] The speech existence coefficient is proportional to the ratio of the minimum power spectrum signal of the channel corresponding to the current speech frame.
[0068] As shown in Figure 6, in one embodiment, the omnidirectional listening speech enhancement method for hearing aids based on microphone array provided by the present invention further includes the following steps in step S130: Step S133, binarizing the speech presence coefficient by setting a threshold, so that when the speech presence coefficient exceeds the set threshold, it is determined that there is target speech, otherwise there is no target speech, thus obtaining a binarized speech representation.
[0069] Step S134: Introduce a probability smoothing factor to smooth the binarized speech representation, and transform the binarized speech representation into the speech existence probability.
[0070] Step S135: Based on the probability of speech presence, a noise figure smoothing factor is introduced to estimate the environmental noise of the power of each channel in order to output the noise power spectrum.
[0071] As shown in Figure 7, in one embodiment, the omnidirectional listening speech enhancement method for hearing aids based on a microphone array provided by the present invention includes the following steps in step S140: Step S141, based on the noise power spectra of the positive narrow beam signal and the surrounding wide beam signal, calculate the power spectra of the target speech signal corresponding to the positive narrow beam signal and the surrounding wide beam signal respectively.
[0072] Step S142: Substitute the power spectrum and noise power of the target speech signal into the Wiener filtering technique, and introduce a smoothing coefficient and a signal-to-noise ratio factor to determine the noise reduction coefficient of the Wiener filter.
[0073] As shown in Figure 8, in one embodiment, the microphone array-based omnidirectional listening speech enhancement method for hearing aids provided by the present invention further includes the following steps in step S140: Step S143, based on the Wiener filter noise reduction coefficients corresponding to the positive narrow beam signal and the surrounding wide beam signal respectively, environmental noise cancellation processing is performed on the positive narrow beam signal and the surrounding wide beam signal to obtain the noise-cancelled positive narrow beam signal and the surrounding wide beam signal.
[0074] Step S144: The positive narrow-beam signal after noise cancellation and the surrounding wide-beam signal are synthesized to output the amplitude spectrum of the target speech signal.
[0075] The following describes the microphone array-based omnidirectional hearing aid speech enhancement system provided by the present invention. The microphone array-based omnidirectional hearing aid speech enhancement system described below can be referred to in correspondence with the microphone array-based omnidirectional hearing aid speech enhancement method described above.
[0076] As shown in Figure 9, in one embodiment, a microphone array-based omnidirectional hearing aid speech enhancement system includes a function construction module 910, a signal acquisition module 920, a noise estimation module 930, and a noise cancellation module 940.
[0077] The function construction module 910 is used to determine the front and rear wide-beam signals respectively based on the frequency domain characterization of the front and rear microphones, and to construct the front and rear amplitude response functions respectively based on the front and rear wide-beam signals.
[0078] The signal acquisition module 920 is used to determine the time-frequency energy distribution of the microphone-acquired signal based on the ratio of the forward and backward amplitude response functions, so as to obtain the forward narrow-beam signal and the surrounding wide-beam signal.
[0079] The noise estimation module 930 is used to estimate the ambient noise of the positive narrow beam signal and the surrounding wide beam signal by using the minimum power spectrum signal within a set time window, in order to determine whether the target speech exists, and outputs the noise power spectrum.
[0080] The noise cancellation module 940 is used to estimate the power spectrum of the target speech signal based on the noise power spectrum, and to call Wiener filtering technology to perform noise cancellation, thereby obtaining the amplitude spectrum of the synthesized target speech signal after noise reduction.
[0081] In this embodiment, the function construction module 910 of the microphone array-based omnidirectional listening speech enhancement system for hearing aids provided by the present invention is specifically used to: obtain the microphone spacing and the signal sampling rate and discrete frequency points of the microphone acquisition signal, and construct the frequency domain representation of the front microphone and the rear microphone based on the microphone spacing, signal sampling rate and discrete frequency points, combined with the speed of sound propagation in the air and the angle between the target sound source and the line connecting the front and rear microphones.
[0082] Based on the frequency domain representations of the forward and backward microphones, forward amplitude response functions and backward amplitude response functions are constructed respectively, and the ratio between the forward amplitude response function and the backward amplitude response function is calculated.
[0083] In this embodiment, the signal acquisition module 920 of the microphone array-based omnidirectional listening voice enhancement system for hearing aids provided by the present invention is specifically used to: determine the degree of deviation of the direction of the sound source in any time-frequency unit of the microphone-acquired signal according to the ratio of the forward and backward amplitude response functions, so as to determine the time-frequency energy distribution. The degree of deviation of the direction of the sound source is proportional to the ratio of the forward and backward amplitude response functions.
[0084] Based on the set forward target angle and the ratio of the forward and backward amplitude response functions, the time-frequency masking matrices of the forward narrow beam signal and the surrounding wide beam signal are obtained respectively. The forward narrow beam signal and the surrounding wide beam signal are then generated based on the time-frequency masking matrices and the frequency domain representation of the forward microphone.
[0085] In this embodiment, the noise estimation module 930 of the microphone array-based omnidirectional listening voice enhancement system for hearing aids provided by the present invention is specifically used to: acquire the power spectrum signals of the positive narrow beam signal and the surrounding wide beam signal, and determine the minimum power spectrum signal within a set time window.
[0086] Calculate the ratio of the power of each channel in the current speech frame to the minimum power spectrum signal, and define the speech presence coefficient of the target speech in each channel of the current speech frame based on the magnitude of the ratio.
[0087] The speech existence coefficient is proportional to the ratio of the minimum power spectrum signal of the channel corresponding to the current speech frame.
[0088] In this embodiment, the noise estimation module 930 of the microphone array-based omnidirectional listening speech enhancement system for hearing aids provided by the present invention is further used to: perform binarization processing on the speech presence coefficient by setting a threshold, so as to determine that the target speech exists when the speech presence coefficient exceeds the set threshold, otherwise the target speech does not exist, thereby obtaining a binarized speech representation.
[0089] A probability smoothing factor is introduced to smooth the binarized speech representation, transforming the binarized speech representation into the probability of speech presence.
[0090] Based on the probability of speech presence, a noise figure smoothing factor is introduced to estimate the environmental noise of each channel power in order to output the noise power spectrum.
[0091] In this embodiment, the noise cancellation module 940 of the microphone array-based omnidirectional listening speech enhancement system for hearing aids provided by the present invention is specifically used to: calculate the power spectrum of the target speech signal corresponding to the positive narrow beam signal and the surrounding wide beam signal based on the noise power spectrum of the positive narrow beam signal and the surrounding wide beam signal, respectively.
[0092] The Wiener filter noise reduction coefficients are determined by substituting the power spectrum and noise power of the target speech signal into the Wiener filtering technique and introducing a smoothing coefficient and a signal-to-noise ratio factor.
[0093] In this embodiment, the noise cancellation module 940 of the microphone array-based omnidirectional listening voice enhancement system for hearing aids provided by the present invention is further used to: perform environmental noise cancellation processing on the positive narrow beam signal and the surrounding wide beam signal based on the Wiener filter noise reduction coefficients corresponding to the positive narrow beam signal and the surrounding wide beam signal respectively, to obtain the noise-cancelled positive narrow beam signal and the surrounding wide beam signal.
[0094] The positive narrow-beam signal after noise cancellation and the surrounding wide-beam signal are synthesized to output the amplitude spectrum of the target speech signal.
[0095] Figure 10 illustrates a schematic diagram of the physical structure of an electronic device, which can be a smart terminal. Its internal structure is shown in Figure 10. The electronic device includes a processor, internal memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used for communication with external terminals via a network connection. When executed by a processor, this computer program implements a microphone array-based omnidirectional hearing aid speech enhancement method. The method includes: determining forward and backward wide-beam signals based on the frequency domain representations of the forward and backward microphones, and constructing forward and backward amplitude response functions based on these signals; determining the time-frequency energy distribution of the microphone-acquired signals based on the ratio of the forward and backward amplitude response functions to obtain a forward narrow-beam signal and surrounding wide-beam signals; estimating the environmental noise of the forward narrow-beam signal and surrounding wide-beam signals by setting the minimum power spectrum signal within a time window to determine the presence of target speech and outputting the noise power spectrum; estimating the power spectrum of the target speech signal based on the noise power spectrum and applying Wiener filtering technology for noise reduction to obtain the amplitude spectrum of the synthesized target speech signal after noise reduction.
[0096] Those skilled in the art will understand that the structure shown in FIG10 is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the electronic device to which the present invention is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0097] On the other hand, the present invention also provides a computer storage medium storing a computer program. When the computer program is executed by a processor, it implements a microphone array-based omnidirectional hearing aid speech enhancement method. The method includes: determining forward and backward wide-beam signals based on the frequency domain characteristics of the forward and backward microphones, and constructing forward and backward amplitude response functions based on the forward and backward wide-beam signals respectively; determining the time-frequency energy distribution of the microphone-acquired signal based on the ratio of the forward and backward amplitude response functions to obtain a forward narrow-beam signal and surrounding wide-beam signals; estimating the environmental noise of the forward narrow-beam signal and surrounding wide-beam signals by setting the minimum power spectrum signal within a set time window to determine whether the target speech exists, and outputting the noise power spectrum; estimating the power spectrum of the target speech signal based on the noise power spectrum, and calling Wiener filtering technology for noise cancellation to obtain the amplitude spectrum of the synthesized target speech signal after noise reduction.
[0098] On another front, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and when executing the computer instructions, implements a microphone array-based omnidirectional hearing aid speech enhancement method. This method includes: determining forward and backward wide-beam signals based on the frequency domain representations of the forward and backward microphones, and constructing forward and backward amplitude response functions based on the forward and backward wide-beam signals respectively; determining the time-frequency energy distribution of the microphone-acquired signal based on the ratio of the forward and backward amplitude response functions to obtain a forward narrow-beam signal and surrounding wide-beam signals; estimating the environmental noise of the forward narrow-beam signal and surrounding wide-beam signals by using a minimum power spectrum signal within a set time window to determine the presence of target speech, and outputting a noise power spectrum; estimating the power spectrum of the target speech signal based on the noise power spectrum, and applying Wiener filtering technology for noise cancellation to obtain the amplitude spectrum of the synthesized target speech signal after noise reduction.
[0099] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory.
[0100] By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0101] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0102] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A method for enhancing omnidirectional listening speech in hearing aids based on a microphone array, characterized in that, The method includes: determining forward and backward wide-beam signals based on the frequency domain characterization of the forward and backward microphones, and constructing forward and backward amplitude response functions based on the forward and backward wide-beam signals; determining the time-frequency energy distribution of the microphone-acquired signal based on the ratio of the forward and backward amplitude response functions to obtain a forward narrow-beam signal and surrounding wide-beam signals; estimating the environmental noise of the forward narrow-beam signal and surrounding wide-beam signals by using the minimum power spectrum signal within a set time window to determine whether the target speech exists, and outputting the noise power spectrum; estimating the power spectrum of the target speech signal based on the noise power spectrum, and using Wiener filtering technology to perform noise cancellation to obtain the amplitude spectrum of the synthesized target speech signal after noise reduction.
2. The method for enhancing omnidirectional hearing aid speech based on a microphone array according to claim 1, characterized in that, The step of determining the front and rear wide-beam signals based on the frequency domain representations of the front and rear microphones, and constructing front and rear amplitude response functions based on the front and rear wide-beam signals, includes: obtaining the microphone spacing and the signal sampling rate and discrete frequency points of the microphone acquisition signals; constructing the frequency domain representations of the front and rear microphones based on the microphone spacing, signal sampling rate, and discrete frequency points, combined with the speed of sound in the air and the angle between the target sound source and the line connecting the front and rear microphones; constructing the front amplitude response function and the rear amplitude response function based on the frequency domain representations of the front and rear microphones, and calculating the ratio between the front amplitude response function and the rear amplitude response function.
3. The method for enhancing omnidirectional hearing aid speech based on a microphone array according to claim 2, characterized in that, The step of determining the time-frequency energy distribution of the microphone-acquired signal based on the ratio of the forward and backward amplitude response functions to obtain a positive narrow-beam signal and a surrounding wide-beam signal includes: determining the degree of deviation of the sound source arrival direction of the microphone-acquired signal within any time-frequency unit based on the magnitude of the ratio of the forward and backward amplitude response functions to determine the time-frequency energy distribution, wherein the degree of deviation of the sound source arrival direction is proportional to the ratio of the forward and backward amplitude response functions; obtaining the time-frequency masking matrices of the positive narrow-beam signal and the surrounding wide-beam signal based on a set positive target angle and the ratio of the forward and backward amplitude response functions, and generating the positive narrow-beam signal and the surrounding wide-beam signal based on the time-frequency masking matrices and the frequency domain representation of the forward microphone.
4. The method for enhancing omnidirectional hearing aid speech based on a microphone array according to claim 1, characterized in that, The step of estimating the environmental noise of the positive narrow-beam signal and the surrounding wide-beam signal by using the minimum power spectrum signal within a set time window to determine whether the target speech exists and outputting the noise power spectrum includes: acquiring the power spectrum signals of the positive narrow-beam signal and the surrounding wide-beam signal, and determining the minimum power spectrum signal within the set time window; calculating the ratio of the power of each channel of the current speech frame to the minimum power spectrum signal, so as to define the speech presence coefficient of the target speech in each channel of the current speech frame according to the size of the ratio; wherein, the speech presence coefficient is proportional to the size of the ratio of the power of the corresponding channel of the current speech frame to the minimum power spectrum signal.
5. The method for enhancing omnidirectional hearing aid speech based on a microphone array according to claim 4, characterized in that, The step of estimating the environmental noise of the positive narrow-beam signal and the surrounding wide-beam signal by using the minimum power spectrum signal within a set time window to determine whether the target speech exists and outputting the noise power spectrum further includes: binarizing the speech presence coefficient by setting a threshold, so that the presence of the target speech is determined when the speech presence coefficient exceeds the set threshold, otherwise the target speech does not exist, thus obtaining a binarized speech representation; introducing a probability smoothing factor to smooth the binarized speech representation, converting the binarized speech representation into a speech presence probability; and based on the speech presence probability, introducing a noise coefficient smoothing factor to estimate the environmental noise of the power of each channel, thus outputting the noise power spectrum.
6. The method for enhancing omnidirectional hearing aid speech based on a microphone array according to claim 1, characterized in that, The step of estimating the target speech signal power spectrum based on the noise power spectrum and applying Wiener filtering technology for noise removal to obtain the amplitude spectrum of the synthesized target speech signal after noise reduction includes: calculating the target speech signal power spectrum corresponding to the positive narrow-beam signal and the surrounding wide-beam signal based on the noise power spectra of the positive narrow-beam signal and the surrounding wide-beam signal respectively; substituting the target speech signal power spectrum and the noise power into the Wiener filtering technology, and introducing a smoothing coefficient and a signal-to-noise ratio factor to determine the Wiener filter noise reduction coefficient.
7. The method for enhancing omnidirectional hearing aid speech based on a microphone array according to claim 1, characterized in that, The step of estimating the power spectrum of the target speech signal based on the noise power spectrum and applying Wiener filtering technology for noise cancellation to obtain the amplitude spectrum of the synthesized target speech signal after noise reduction further includes: performing environmental noise cancellation processing on the positive narrow-beam signal and the surrounding wide-beam signal based on the Wiener filter noise reduction coefficients corresponding to the positive narrow-beam signal and the surrounding wide-beam signal respectively, to obtain the noise-cancelled positive narrow-beam signal and the surrounding wide-beam signal; and synthesizing the noise-cancelled positive narrow-beam signal and the surrounding wide-beam signal to output the amplitude spectrum of the target speech signal.
8. A microphone array-based omnidirectional hearing enhancement system for hearing aids, characterized in that, To implement the microphone array-based omnidirectional listening speech enhancement method for hearing aids according to any one of claims 1 to 7, the system comprises: a function construction module, configured to determine forward and backward wide-beam signals respectively based on the frequency domain characterization of the forward and backward microphones, and construct forward and backward amplitude response functions respectively based on the forward and backward wide-beam signals; a signal acquisition module, configured to determine the time-frequency energy distribution of the microphone-acquired signal based on the ratio of the forward and backward amplitude response functions, so as to acquire a forward narrow-beam signal and a surrounding wide-beam signal; a noise estimation module, configured to estimate the environmental noise of the forward narrow-beam signal and the surrounding wide-beam signal by using the minimum power spectrum signal within a set time window, so as to determine whether the target speech exists, and output the noise power spectrum; and a noise cancellation module, configured to estimate the power spectrum of the target speech signal based on the noise power spectrum, and call Wiener filtering technology to perform noise cancellation, so as to obtain the amplitude spectrum of the synthesized target speech signal after noise reduction.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the microphone array-based omnidirectional listening speech enhancement method for hearing aids as described in any one of claims 1 to 7.
10. A computer storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the microphone array-based omnidirectional listening speech enhancement method for hearing aids as described in any one of claims 1 to 7.