A method for sound source localization and enhancement of a binaural hearing aid

By combining random maximum likelihood estimation and Gaussian hybrid model to optimize sound source positioning and combined with generalized side lobe phase destroyers for adaptive noise cancellation, the problems of inaccuracy and insufficient noise suppression of binaural hearing aids in complex acoustic environments are solved, and higher positioning accuracy and signal enhancement effects are achieved.

CN119789033BActive Publication Date: 2025-07-11SHENZHEN PURETING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510286759.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-11
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

In complex acoustic environments, existing binaural hearing aids have insufficient accuracy and robustness in sound source positioning, especially in multi-sound sources and strong interference, and the sound source signal intensity is reduced.

Method used

Using a sound source positioning method combining random maximum likelihood estimation and Gaussian hybrid model, the sound source identification and tracking are optimized by analyzing the covariance matrix of the received signal of the microphone and the head-related transfer function, and adaptive noise cancellation is performed using a generalized side lobe erase generator.

Benefits of technology

Improves the accuracy of sound source positioning and noise cancellation performance, can adapt to different users and environment changes in real time, and provide an improved auditory experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119789033B_ABST
    Figure CN119789033B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of sound source localization for hearing aids, and specifically relates to a sound source localization and enhancement method for binaural hearing aids. The method includes: determining the sound source angle at any time period according to the correlation between the received signals of different microphones at any time period and combining random maximum likelihood estimation; determining the existence probability of each sound source at the sound source angle at any time period by analyzing the distribution of each sound source and combining the Gaussian mixture model; determining the dynamic existence probability of each sound source at the sound source angle at any time period based on the existence probability of each sound source at the sound source angle at any time period; enhancing the sound source based on the dynamic existence probability of each sound source at the sound source angle of the current time period. The purpose of this application is to improve the accuracy of sound source localization and enhance the intensity of the sound source signal received by the hearing aid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of hearing aid sound source localization, and particularly relates to a sound source localization and enhancement method for binaural hearing aids. Background Art

[0002] With the aging of society and the increase in the number of people with hearing impairments, the demand for hearing aids is constantly growing. Hearing aids (HAs) are auxiliary devices used to help hearing-impaired people improve their hearing. They improve the user's auditory experience by amplifying sound. Most modern hearing aids use multi-microphone arrays and advanced signal processing algorithms to improve the clarity of speech and suppress background noise. These hearing aids utilize beamforming technology to enhance the sound from a specific direction through spatial filtering while suppressing the sound from other directions.

[0003] In the prior art, during the improvement and application of various beamforming algorithms to binaural hearing aids, in the face of complex acoustic environments, such as dealing with multiple sound sources, strong interference, reverberation and other complex situations, the accuracy and robustness of sound source localization are insufficient; algorithms such as SML and AML are used to reduce the computational complexity of the maximum likelihood estimation method, but it is not specifically designed for the hearing aid application scenario and does not consider the head-related transfer function. When applied to binaural hearing aids, due to the influence of head rotation, the accuracy of sound source localization is reduced, and at the same time, the intensity of the sound source signal received by the hearing aid is reduced. Summary of the Invention

[0004] To solve the above technical problems, this application provides a sound source localization and enhancement method for binaural hearing aids to solve the existing problems.

[0005] A sound source localization and enhancement method for binaural hearing aids in this application adopts the following technical solutions:

[0006] An embodiment of this application provides a sound source localization and enhancement method for binaural hearing aids, and this method includes the following steps:

[0007] When there are multiple sound sources in the environment, obtain the received signals of each microphone within a preset time period, and divide the preset time period into multiple time segments;

[0008] According to the correlation between the received signals of different microphones in any time segment, and in combination with the random maximum likelihood estimation, determine the sound source angle in any time segment;

[0009] At the sound source angle of any time segment, by analyzing the distribution of each sound source, and in combination with the Gaussian mixture model, determine the existence probability of each sound source at the sound source angle of any time segment;

[0010] Determine the dynamic presence probabilities of each sound source at the sound source angle for any time period based on the presence probabilities of each sound source at the sound source angle for any time period.

[0011] Based on the received signals of each microphone in the current time period, obtain the beam in the current time period; based on the dynamic presence probabilities of each sound source at the sound source angle in the current time period, determine the iteration step size of each sound source at the sound source angle in the current time period, and combine the beam in the current time period and the received signals of each microphone to determine the adaptive noise cancellation coefficient of each sound source in the next time period, and enhance the sound source.

[0012] Preferably, the method for determining the sound source angle for any time period is as follows:

[0013] Take the set composed of the received signals of all microphones in any time period as the microphone signal, obtain the covariance matrix of the microphone signal, use it as the input of the stochastic maximum likelihood estimation algorithm, iterate the stochastic maximum likelihood estimation algorithm, and output the sound source angle in any time period.

[0014] Preferably, the method for determining the presence probability of each sound source at the sound source angle for any time period is as follows:

[0015] The expression of the presence probability of sound source q at the sound source angle in time period b is: ; In the formula, represents the prior probability of the q-th Gaussian component in the Gaussian mixture model, where one Gaussian component represents one sound source; represents the q-th Gaussian component; represents the sound source angle in time period b; , respectively represent the means of the q-th Gaussian component and the i-th Gaussian component in the Gaussian mixture model; , respectively represent the variances of the q-th Gaussian component and the i-th Gaussian component in the Gaussian mixture model; represents the Gaussian distribution function; Q represents the number of all Gaussian components.

[0016] Preferably, the prior probability, mean and variance of the Gaussian component are all estimated by the expectation maximization algorithm.

[0017] Preferably, the method for determining the dynamic presence probability of each sound source is as follows:

[0018] Take the presence probabilities of each sound source at the sound source angle for any time period as the input of the speech separation SPP algorithm, and output the dynamic presence probabilities of each sound source at the sound source angle for any time period.

[0019] Preferably, the expression of the iteration step size of each sound source at the sound source angle in the current time period is: ; where represents the iteration step of sound source q at the sound source angle in the current time period; represents the existence probability of sound source q at the sound source angle in the current time period; represents a preset value.

[0020] Preferably, the process of obtaining the beam in the current time period is as follows:

[0021] Taking the microphone signal in the current time period as the input of the minimum variance distortionless response algorithm, and outputting the beam in the current time period.

[0022] Preferably, determining the adaptive noise cancellation coefficient of each sound source in the next time period includes:

[0023] The adaptive noise cancellation coefficient of sound source q in the next time period has the following expression: ; where represents the adaptive noise cancellation coefficient of sound source q in the current time period, where the initial value of is 0; represents the beam in the current time period; represents the adaptive blocking matrix coefficient of sound source q in the next time period; represents the microphone signal in the current time period; represents taking the absolute value.

[0024] Preferably, the adaptive blocking matrix coefficient of sound source q in the next time period is the output result obtained by taking the microphone signal in the next time period and the dynamic existence probability of sound source q as the input of the generalized sidelobe canceller.

[0025] Preferably, enhancing the sound source includes:

[0026] Taking the microphone signal in the next time period as the input of the minimum distortionless variance, outputting the preliminarily enhanced signal, and taking the enhanced signal as the input of the generalized sidelobe cancellation technique. Among them, taking the adaptive noise cancellation coefficients of each sound source in the next time period as the adaptive filter coefficients in the adaptive filter of the generalized sidelobe canceller, and outputting the finally enhanced signal.

[0027] This application has at least the following beneficial effects:

[0028] This application optimizes the identification and tracking of sound sources by integrating beamforming technology and statistical models; optimizes the Gaussian mixture model (GMM) by introducing multi-channel speech presence probability (SPP) estimation, further improving the accuracy of sound source localization. Further, by adapting the blocking matrix and noise canceller of the generalized sidelobe canceller CGC, the performance of noise cancellation is enhanced; this application maintains high performance even when the user's head rotates by considering the head-related function; the advantage of this application is that it can process audio signals in real time, adapt to different users without prior training, making it have good real-time performance and versatility. This application can provide an improved auditory experience under various noise conditions, improve the accuracy of sound source localization and enhance the intensity of the sound source signal received by the hearing aid. Description of the Drawings

[0029] To more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0030] Figure 1 It is a flowchart of the steps of a sound source localization and enhancement method for a binaural hearing aid provided by an embodiment of the present application;

[0031] Figure 2 It is a flowchart of sound source localization provided by an embodiment of the present application. Detailed Embodiments

[0032] To further elaborate on the technical means and effects adopted by the present application to achieve the intended invention purpose, the following, in combination with the drawings and preferred embodiments, details the specific embodiments, structures, features and effects of a sound source localization and enhancement method for a binaural hearing aid proposed according to the present application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.

[0034] The following specifically describes the specific solution of a sound source localization and enhancement method for a binaural hearing aid provided by the present application in conjunction with the drawings.

[0035] An embodiment of the present application provides a sound source localization and enhancement method for binaural hearing aids. Specifically, the following sound source localization and enhancement method for binaural hearing aids is provided. Please refer to Figure 1 , and the method includes the following steps:

[0036] Step S1: Obtain the received signals of each microphone within a preset duration, and divide the preset duration into multiple time periods. In this embodiment, the binaural hearing aids use a microphone array to capture sounds from different directions, which includes at least four microphones. The microphone array can be any one of the following: a one-dimensional linear array, a two-dimensional planar array, and a three-dimensional stereo array. Obtain the received signals of each microphone within a preset duration, and divide the preset duration into multiple time periods.

[0037] Preferably, two microphones are installed on a single hearing aid, and the distance between them is 15.32 mm. The installation position is at the ear. The received signal is , where n represents the sampling moment, T represents the transpose, and the frequency domain representation is , where f represents the frequency point. The sampling frequency of the microphone is between 8 kHz and 48 kHz, and the collected signals are segmented, that is, frame processed, and the duration of each frame is 20 ms.

[0038] In addition, it should be noted that the value of the preset duration is set artificially. In this embodiment, the value of the preset duration is 1 s, and the implementer can also set it according to the specific situation. This embodiment does not make special restrictions.

[0039] Step S2: Determine the sound source angle in any time period according to the correlation between the received signals of different microphones in any time period and combined with the random maximum likelihood estimation.

[0040] The traditional generalized sidelobe canceller (GSC) usually processes the signals received by the microphone array based on a preset direction pattern or algorithm, such as the minimum variance distortionless response (MVDR). Its input is the original audio signal containing multiple sound sources and noise captured by the microphone array. By performing a weighted summation operation on these signals, a main beam is formed in the desired direction, so that there is a relatively high gain in the direction of the target sound source, but it may still contain certain interference and noise. The main beam is the sound source direction, but its sound source localization depends more on the fixed beam direction and noise. The main beam is the sound source direction, but its sound source depends more on the fixed beam direction and simple signal difference judgment. In a complex multi-source environment, it is difficult to accurately identify the direction of each sound source, and it is easy to have positioning deviation, which affects the subsequent signal processing effect.

[0041] In this embodiment, the Stochastic Maximum Likelihood (SML) estimation function is used to preliminarily estimate the direction of the sound source. Using the stochastic maximum likelihood method, considering the Head-Related Transfer Function (HRTF) and the position noise conditions, the direction of the sound source is estimated by constructing and maximizing the log-likelihood function, which can more accurately locate the sound source in a complex acoustic environment, provide accurate sound source position information for subsequent signals, and improve the positioning accuracy and stability of the entire system.

[0042] The traditional Stochastic Maximum Likelihood (SML) estimation is the logarithmic function of the probability density function of the multivariate Gaussian distribution. In this embodiment, however, the log-likelihood function is constructed based on the spatial covariance matrix of the signal and the spatial covariance matrix of the noise. From the perspective of information theory, the log-likelihood function is the logarithmic form that describes the probability of the observed data given the model parameters. For the signals received by the microphone array, it is desired to find a parameter (the direction of the sound source) such that the probability of the observed data X(n) occurring under this parameter is maximized. First, estimate the powers of the clean signal and the noise signal based on the sound signals received by the microphones, and then use these estimated values to optimize the log-likelihood function to extract information related to the direction of the sound source.

[0043] Therefore, according to the correlation between the received signals of different microphones at any time period and combined with the stochastic maximum likelihood estimation, the angle of the sound source at any time period is determined. Specifically:

[0044] Take the set of the received signals of all microphones at any time period as the microphone signal, obtain the covariance matrix of the microphone signal as the input of the stochastic maximum likelihood estimation algorithm, iterate the stochastic maximum likelihood estimation algorithm, and output the angle of the sound source at any time period. To more specifically illustrate the principle of obtaining the sound source angle, the specific principle derivation process of the sound source angle is as follows:

[0045] The microphone signal includes the received sound source signal and the noise signal, where the received sound source is denoted as the target signal. Assume that the target signal is an unknown deterministic signal, and the noise signal follows a zero-mean complex Gaussian distribution. First, calculate the probability density function of the microphone received signal through the covariance matrix, take the logarithm of it to obtain the log-likelihood function, and then integrate the data of B frames. Then, for the log-likelihood function of a single sound source signal:

[0046] (1)

[0047] In the formula, B represents the number of all time periods within the preset duration; M represents the number of microphones; Tr{} represents taking the trace of the matrix; represents the direction of the sound source, which is the parameter to be estimated, and its optimal value is determined by maximizing the log-likelihood function.

[0048] represents the covariance matrix of the microphone signal, and the calculation formula is:

[0049] (2)

[0050] In the formula, represents the head-related transfer functions (HRTFs) related to the sound source direction and frequency k, considering the influence of head movement; represents the power spectral value of the target signal in the microphone signal at frequency k in the frequency domain; represents the power spectrum of the noise signal in the microphone signal; represents the covariance matrix of the noise signal; represents the conjugate transpose of

[0051] Among them, the process of obtaining the covariance matrix of the zero-mean complex Gaussian distributed noise signal is a well-known technology, and the head-related transfer function is a well-known technology, and their specific principles will not be elaborated here.

[0052] The head-related transfer function described above is a function that describes the frequency response changes caused by the influence of human physiological structures such as the head, pinna, and torso when sound propagates from the sound source to both ears. It plays a crucial role in the field of acoustic signal processing. When sound propagates in space and reaches both ears, the head will cause phenomena such as occlusion, reflection, and diffraction of sound, and the unique shape of the pinna will further modulate the sound signal.

[0053] These factors cause changes in the amplitude and phase of sound at different frequencies, and the head-related transfer function HRTFs can accurately quantify these changes. In this embodiment, by considering the head-related transfer function in multiple links from parameters to signal processing, the system can take into account the influence of head-related factors on sound propagation, so as to perform sound source localization and signal processing more accurately. For example, in a complex acoustic environment, the sounds emitted by sound sources in different directions are modulated by the head-related transfer function, and the signals received at both ears will show different characteristics. The system can better distinguish the direction and distance of the sound source by analyzing these characteristic differences, improve the performance of the hearing aid in a multi-source and multi-complex noise environment, and provide a more accurate sound enhancement and a clearer auditory experience for users.

[0054] The covariance matrix of the microphone signal consists of two parts. One part is the covariance of the target signal modulated by the head-related transfer function , and the other part is the noise covariance , 。The head-related transfer function takes into account the influence of the head, such as occlusion and reflection, on sound during the propagation of sound from the sound source to the microphone. It modulates the signal source at different frequencies and directions. The power spectrum of the target signal reflects the energy distribution of the target signal at different frequencies, the power spectrum of the noise signal represents the average power of the noise at each frequency, and the noise correlation matrix describes the correlation of the noise between different microphones.

[0055] Among them, the power spectrum estimates of the target signal and the noise signal are as follows:

[0056] (3)

[0057] (4)

[0058] In the formula (3) for the power spectrum estimate of the target signal, is the covariance matrix of the microphone signals, which contains the mixed information of the target signal and the noise signal; is the head-related transfer function, which represents the linear transformation relationship during the propagation of sound from the sound source to the microphone. This transformation reflects the characteristic changes of sound due to the influence of the head at different frequencies and directions; is equivalent to initially removing the influence of noise from the mixed covariance information, which is an operation of feature screening; then through and a matrix multiplication operation is performed, which is a combination of the conjugate transpose of the head-related transfer function and the inverse of the noise correlation matrix to perform a further linear transformation on the preliminarily processed signal to highlight the characteristics of the target signal. Among them, represents the inverse matrix of the covariance matrix of the noise signal; finally, normalization is performed through the operation of the denominator, so that the obtained can accurately represent the power spectrum characteristics of the target signal at different frequencies, that is, the purpose of extracting the power spectrum of the target signal from the mixed signal is achieved.

[0059] In the formula (4) for the power spectrum estimate of the noise signal, Tr{} represents the trace of the matrix; represents the identity matrix; involves taking the inverse of the matrix related to the head-related transfer function and the noise correlation matrix to adjust the signal space for subsequent extraction of the characteristics of the noise signal; through matrix multiplication operations with matrices such as and the trace operation of the matrix, the adjusted signal is comprehensively processed to extract the power spectrum information of the noise signal from the overall signal and matrix relationship. The trace operation of the matrix is equivalent to summing or extracting certain characteristics of the matrix after a series of transformations, and finally the power spectrum of the noise signal is obtained, achieving the goal of accurately estimating the noise power spectrum from the complex signal and matrix combination.

[0060] First, calculate the initial covariance matrix of the signal using the traditional method. Obtain the estimated power spectrum of the target signal according to formula (3), obtain the estimated power spectrum of the noise signal according to formula (4), then update the covariance through the covariance matrix formula (2), and then gradually adjust the covariance matrix, the estimated values of the target signal power spectrum and the noise signal power spectrum through iteration, making them closer and closer to the true values.

[0061] Substitute the finally estimated noise power and signal power into the log-likelihood function formula (1) to obtain the objective function (5):

[0062] (5)

[0063] where and are two auxiliary matrices used for projecting and processing the target signal and the noise. The calculation is obtained from formulas (6) and (7), specifically:

[0064] (6)

[0065] (7)

[0066] The construction process of is actually a recombination and adjustment operation of the signal and the noise after being modulated by HRTFs. By first calculating and taking the inverse, and then multiplying with and

[0067] The purpose is to highlight the part of the target signal in the construction and subsequent calculation of the log-likelihood function, so that in a complex acoustic environment, the information related to the target sound source can be more effectively separated from the mixed signal. For it is the complementary space matrix of and its function is to shield the target signal. In the calculation of the log-likelihood function and the entire signal processing flow, when is used to extract the target signal,

[0068] can process and isolate other parts unrelated to the target signal (mainly noise and interference signals). The two cooperate with each other to make the entire signal processing process more accurate and efficient, thereby improving the accuracy of sound source localization and signal enhancement. (8)

[0069] In the formula, represents the sound source angle in time period b; b represents time period b. represents the number of all frequency points of time period b in the frequency domain; argmax( ) represents the function for finding the maximum index.

[0070] Compared with the traditional Stochastic Maximum Likelihood (SML) estimation method that directly locates the sound source according to the probability density function of the simple signal model, this embodiment not only considers the covariance matrices of the target signal and the noise signal in the calculation of the log-likelihood function, but also takes into account the Head-Related Transfer Function (HRTF) to modulate the target signal, so as to better adapt to the complex acoustic environment and improve the accuracy and robustness of sound source localization.

[0071] Among them, the Stochastic Maximum Likelihood (SML) estimation method is a well-known technology, and its specific principle will not be elaborated here.

[0072] Thus, the sound source angle is obtained.

[0073] Step S3: At the sound source angle of any time period, by analyzing the distribution of each sound source and combining with the Gaussian Mixture Model (GMM), determine the existence probability of each sound source at the sound source angle of any time period.

[0074] The Stochastic Maximum Likelihood (SML) estimation provides the preliminary direction information of the sound source. However, in the real acoustic environment, there are multiple sound sources and the background noise is complex and variable. To improve the accuracy of localization, it is necessary to statistically analyze the estimated narrowband angles. In this embodiment, the Gaussian Mixture Model (GMM) is used to calculate the existence probability of the target signal.

[0075] The Gaussian Mixture Model (GMM) can describe such a complex probability distribution by combining multiple Gaussian components. Each Gaussian component can represent a potential sound source or noise source, and its mean and variance can capture the characteristics of the sound source or noise in the spatial frequency. Moreover, the Gaussian Mixture Model (GMM) can adapt to this change by adjusting the parameters of the Gaussian components (such as prior probability, mean, and variance). When a new sound source appears or an existing sound source disappears, the Gaussian Mixture Model (GMM) can re-estimate the weights and parameters of each Gaussian component to reflect the change of the acoustic environment. This adaptive ability is crucial for devices such as hearing aids because it needs to always maintain good performance in different usage scenarios and accurately identify and track the target sound source.

[0076] Assume there are Q sound sources. Then, for each sound source, the method for determining the existence probability at the sound source angle of any time period is as follows:

[0077] Expression of the existence probability of sound source q at the sound source angle of time period b is:

[0078] (9)

[0079] In the formula, represents the prior probability of the q-th Gaussian component in the Gaussian mixture model, where one Gaussian component represents one sound source; represents the q-th Gaussian component; represents the sound source angle in time period b; 、 respectively represent the position expectations of the q-th Gaussian component and the i-th Gaussian component in the Gaussian mixture model; 、 respectively represent the position variances of the q-th Gaussian component and the i-th Gaussian component in the Gaussian mixture model; represents the Gaussian distribution function; Q represents the number of all Gaussian components.

[0080] Among them, the prior probability, mean, and variance of the Gaussian component are all estimated by the expectation maximization algorithm.

[0081] It is an empirical estimate based on the general understanding of the sound source distribution in different acoustic environments or obtained through statistical analysis of a large amount of experimental data. When calculating the sound source presence probability, as the initial weight of each Gaussian component, combined with the Gaussian distribution function, it affects the final probability calculation result. It reflects the relative importance of each potential sound source or noise source in the model without specific observation data. As the system receives new audio signals and processes them, its value will be dynamically adjusted according to the actual data during the subsequent expectation maximization (EM) algorithm iteration process to more accurately reflect the actual situation of each sound source in the current acoustic environment.

[0082] is the position expectation of the q-th sound source, represents the position variance of the q-th sound source, which is also set based on the understanding of common acoustic signal characteristics in the initial stage. In terms of speech signals, according to factors such as the frequency range of speech, propagation characteristics in space, and performance in different environments, a rough initial mean and variance are determined. As the system runs, by analyzing and processing the signals received by the microphone array, the EM algorithm is used to statistically analyze the continuous time frame data to continuously update these two parameters. In each iteration, according to the matching degree of the current data point with each Gaussian component (calculated by the likelihood function), their values are adjusted to more accurately reflect the characteristic changes of the actual sound source in space and frequency.

[0083] ( ) represents the Gaussian distribution function, Represents the estimated value of the sound source direction. The entire formula calculates the ratio of the probability contribution of a specific sound source to the total sum of the probability contributions of all sound sources, obtaining the probability of the q-th sound source existing when the current estimated sound source direction is given. Such a calculation method can comprehensively consider multiple potential sound sources or noise sources in a complex acoustic environment, and accurately evaluate the possibility of each sound source existing in each time-frequency frame based on their respective prior probabilities, position characteristics, and the current sound source direction estimate. This provides a crucial basis for subsequent signal processing decisions (such as beamforming, noise suppression, etc.), and helps improve the performance of sound source localization and signal enhancement of hearing aids in complex acoustic scenarios.

[0084] In this embodiment, the Expectation-Maximization (EM) algorithm is used to estimate the mean (representing the sound source direction), variance (indicating the uncertainty of the direction), and mixing coefficient (prior probability) of each Gaussian component. The EM algorithm performs statistics on a continuous time frame. It starts with an initial guess, then calculates the expected value of the hidden variable in the Expectation (E) step, and then re-estimates the model parameters in the Maximization (M) step to maximize the likelihood function of the observed data. This iterative process is continuously alternated until it converges to a set of optimal parameters, enabling the model to accurately reflect the statistical characteristics of the sound source and noise.

[0085] The Expectation-Maximization (EM) algorithm can adapt to changes in the number of sound sources, automatically estimate the number of active sound sources in the environment, and increase the flexibility and adaptability of the system. With these optimized parameters, the beamformer can more effectively separate the target sound source, improve speech clarity, and reduce background noise, significantly enhancing the auditory experience of hearing aid users.

[0086] In traditional sound source localization, mainly methods based on time difference of arrival are used. The sound source position is calculated by measuring the time difference of the sound signal arriving at different microphones, and its principle is based on simple geometric relationships and assumptions about the signal propagation speed. For methods based on beamforming, the signals received by the microphone array are weighted and summed, so that the signals in a specific direction are enhanced, thereby determining the sound source direction, which mainly relies on the spatial domain filtering principle in signal processing.

[0087] Traditional methods can achieve good results in relatively simple and less noisy environments. However, due to the fact that they do not fully consider factors such as the complex probability distribution of acoustic signals and the uncertainty of noise in the actual environment, in complex environments, such as scenarios with multipath propagation, reverberation, and non-stationary noise, their positioning accuracy and reliability will be significantly affected.

[0088] However, the actual acoustic environment often changes dynamically. The number, position, intensity of sound sources, and the characteristics of noise, etc., may all change significantly over time. Traditional sound source localization methods usually rely on static model assumptions and fixed parameter settings, and it is difficult to quickly and effectively adapt to such a dynamically changing environment, resulting in a sharp decline in localization accuracy and reliability when the environment changes.

[0089] Therefore, in this solution, by introducing the presence probability of sound sources and combining the improved Stochastic Maximum Likelihood (SML) estimation, the direction of the sound source is initially determined. By constructing a log-likelihood function and maximizing it to solve, it is possible to estimate the direction of the sound source on the basis of considering complex factors such as Head-Related Transfer Functions (HRTFs) and unknown noise conditions. However, in an actual complex acoustic environment, relying solely on the SML method may not be able to fully and accurately describe the probability distribution of the signal, thus affecting the localization accuracy. The Gaussian Mixture Model (GMM) describes complex probability distributions through the combination of multiple Gaussian components and can more precisely characterize the characteristics of acoustic signals in different frequencies and spaces. By combining the GMM with the SML estimation, the GMM can provide a more accurate signal probability distribution model for the SML estimation. When calculating the log-likelihood function and estimating the direction of the sound source, the SML estimation can better consider the uncertainty of the signal and the complex probability distribution, thus significantly improving the accuracy of sound source localization. In harsh environments with multiple sound sources, strong noise, and complex reflections, etc., it is still possible to relatively accurately determine the position of the sound source.

[0090] Among them, the Gaussian Mixture Model (GMM), the Gaussian distribution function, and the Expectation-Maximization (EM) algorithm are all well-known technologies, and their specific principles will not be elaborated here.

[0091] At this point, the presence probability of the sound source is obtained.

[0092] Step S4: Based on the presence probabilities of each sound source at the sound source angle of any time period, determine the dynamic presence probabilities of each sound source at the sound source angle of any time period.

[0093] In the Gaussian mixture model (GMM), the position of each sound source is represented by a Gaussian component, and its prior probability and variance are usually fixed or set empirically. In a complex and changing acoustic environment, such as the user's head movement, the sudden appearance or disappearance of surrounding sound sources, the change of noise intensity and frequency characteristics, especially when the noise conditions change or the number of sound sources is uncertain, it is difficult to accurately capture the probability of the existence of sound sources in the actual environment. In this embodiment, SPP is used for improvement, specifically: the existence probability of each sound source at the sound source angle in any time period is used as the input of the speech separation SPP algorithm, and the dynamic existence probability of each sound source at the sound source angle in any time period is output. For the sake of easy understanding, the following gives a more specific derivation process, specifically:

[0094] The speech separation SPP algorithm estimates that an additional Gaussian component is introduced in the Gaussian mixture model (GMM) to represent noise. The mean of this noise component is set to the direction of the noise, and the variance is set to the power of the noise. In this way, the GMM can not only model sound sources but also model the noise environment. For each sound source q, two hypotheses are defined:

[0095] Hypothesis that the sound source exists , Hypothesis that the sound source does not exist and only noise exists . Using this hypothesis, the expression for the dynamic existence probability of each sound source at the sound source angle in any time period is:

[0096] (10)

[0097] In the formula, represents the posterior probability to be calculated, that is, the probability that the q-th sound source exists under the condition that the observed data X is known; represents the probability that the observed data X appears under the condition that the sound source q exists, which measures the rationality of the observed data X under the hypothesis that the sound source exists; represents the prior probability that the q-th sound source exists, that is, the probability estimate of the existence of the q-th sound source before the observed data X. This probability is usually obtained based on historical data or domain knowledge; represents the sum of the products of the likelihood function and the prior probability for all hypotheses of the existence of sound sources. Here, all possible sound sources (a total of Q sound sources) are considered, and the probabilities of their existence are weighted and summed; represents the probability that the observed data X appears under the hypothesis that only noise exists, which measures the rationality of the observed data X in the case of no sound source (only noise); represents the prior probability of no speech, that is, the probability of only noise in the case of no existence of any sound source.

[0098] Among them, The prior probability indicating the existence of the q-th sound source, the method for obtaining which is a well-known technique and the specific obtaining process will not be elaborated here. It serves as an important weight factor when calculating the probability of the existence of the sound source. It reflects the likelihood of the existence of sound source q before specific observation data, based on past experience or statistical analysis of a specific environment. In the subsequent construction of the likelihood ratio function, it combines with the probability part calculated based on the observation data to jointly determine the final probability of the existence of sound source q under the given signal X. It is usually obtained through the analysis and statistics of a large amount of audio data. In the application scenario of hearing aids, audio samples may be collected in different environments, and information such as the frequency and location of the sound source appearing in them is statistically analyzed to determine the prior probabilities of different sound sources.

[0099] The prior probability indicating the absence of speech, the obtaining process of which is a well-known technique and the specific principle will not be elaborated here. It corresponds to the prior probability of the existence of the sound source and represents the prior probability in the case of no speech signal (i.e., only noise). In the likelihood ratio function, it is part of the denominator and is used to measure the relative likelihood of the observation data under the hypothesis of no speech, and is compared with the situation under the hypothesis of the existence of the sound source. Through the ratio with and other related calculations, the difference between the two situations of the existence and non-existence of the sound source can be highlighted, so as to more accurately judge whether the sound source exists.

[0100] Define the generalized likelihood ratio, denoted as:

[0101] (11)

[0102] The likelihood ratio is an index for measuring the existence or non-existence of a signal. Then the probability hypothesis can be expressed as the following formula:

[0103] (12)

[0104] Using the original input signal, according to formulas (10), (11), and (12), the likelihood ratio can be expressed as:

[0105] (13)

[0106] (14)

[0107] The likelihood ratio function is constructed based on the covariance matrix of the signal, the noise correlation matrix, and the noise power spectrum. In signal processing, by constructing such a likelihood ratio function, various characteristics of the signal and noise, as well as prior probability information, can be comprehensively considered, and the presence of the sound source can be judged more accurately from a probabilistic perspective. When the likelihood ratio is greater than 1, it tends to indicate the presence of the sound source; when it is less than 1, it tends to indicate the absence of the sound source. It can use a mathematical model to quantitatively judge the presence or absence of the sound source in a complex acoustic environment, overcoming the limitations of simple threshold judgment.

[0108] It is defined as the prior signal-to-noise ratio of the sound source q. Using the matrix inversion lemma, the input signal covariance matrix can be expressed as:

[0109] (15)

[0110] The relevant terms in the likelihood ratio function are further adjusted to make them coordinated with the previous parameters and operation results, more accurately reflecting the mutual relationship between the signal and noise at the covariance matrix level, thereby improving the accuracy of the likelihood ratio function in judging the presence of the sound source.

[0111] Using the above formula, substituting (14) and (15) into the likelihood ratio (13) can be expressed as:

[0112] (16)

[0113] Among them, let , and by calculating the power of the transient MVDR beam:

[0114] (17)

[0115] It reflects the signal energy distribution of the sound source considering the noise correlation and the sound source propagation characteristics. In a multi-source and complex noise environment, this power value can be used as an important indicator to distinguish the intensity and position characteristics of different sound sources. When performing sound source localization, if the power value corresponding to a sound source is large, it indicates that the signal intensity of the sound source is relatively strong in the current environment, and a higher weight will be given in the localization process, which helps to more accurately determine the position of the sound source. At the same time, when calculating the posterior signal-to-noise ratio, it is also a key component. The posterior signal-to-noise ratio is further used in the calculation of the likelihood ratio function, thereby more precisely judging the probability of the presence of the sound source, providing an important basis for the entire signal processing process, and realizing the effective identification and processing of sound sources in a complex acoustic environment.

[0116] represents the channel matrix, which contains the channel gain and phase information from the sound source q to the microphone array. The method of obtaining the channel matrix is a well-known technology, and the specific acquisition process will not be elaborated here. is the conjugate transpose matrix of the channel matrix.

[0117] can be the posterior signal-to-noise ratio of the sound source signal q:

[0118] (18)

[0119] Substitute the prior signal-to-noise ratio (which is obtained based on the pre-estimation of the signal power spectrum of the sound source or based on past experience and statistical analysis in a specific environment) and the posterior signal-to-noise ratio (18) into formula (16), and the likelihood ratio function can be expressed as follows (19):

[0120] (19)

[0121] Compared with the traditional speech separation SPP algorithm that uses a fixed probability density function, using the likelihood ratio function can obtain more accurate sound source localization statistics. It can update the calculation results in real time according to the received signal, so as to adapt to the dynamic changes of the acoustic environment. When new sound sources appear in the environment, the noise intensity or frequency characteristics change, the user's head moves, etc., the statistical characteristics of the signal will change.

[0122] The speech separation SPP algorithm can calculate the dynamic presence probability in real time according to the current acoustic environment and the received signal, and can quickly respond to these changes. By using the statistical characteristics of the signal and noise, considering two hypotheses of the presence and absence of the sound source, and calculating the probability of each hypothesis of the given noise signal in these two cases, it can provide a more dynamic and data-driven probability estimate for the presence of each sound source, making up for the deficiencies of the Gaussian mixture model GMM.

[0123] When calculating the probability of the presence of each time-frequency frame, the speech separation SPP algorithm not only considers the signal characteristics of the current time-frequency frame, but also combines the signal changes of adjacent time-frequency frames. Through the time series analysis of the signal, it can more accurately capture the dynamic change laws of the sound source and noise, so as to optimize the probability calculation model. In addition, SPP also fully considers the uncertainty and complexity of the noise. By analyzing the statistical characteristics of the noise and introducing a more accurate noise model, it can more effectively suppress the interference of the noise in the probability calculation process and improve the accuracy and reliability of the probability estimation of the presence of the sound source.

[0124] By providing the dynamic sound source presence probability, the speech separation SPP algorithm can optimize the performance of the entire adaptive binaural beamformer system. It can more accurately adapt to the actual sound environment, so as to better achieve sound source separation and noise suppression, improve the overall performance of the hearing aid in complex acoustic environments, and meet the hearing assistance needs of users in different scenarios.

[0125] Among them, the voice separation SPP algorithm is a well-known technology, and its specific principle will not be elaborated here.

[0126] Preferably, the flowchart of sound source localization provided in this embodiment is as Figure 2 shown.

[0127] Thus, the dynamic sound source existence probability and the posterior SNR are obtained, improving the accuracy of sound source localization.

[0128] Step S5: Based on the received signals of each microphone in the current time period, obtain the beam of the current time period; based on the dynamic existence probabilities of each sound source at the sound source angle of the current time period, determine the iteration step size of each sound source at the sound source angle of the current time period, and combine the beam of the current time period and the received signals of each microphone to determine the adaptive noise cancellation coefficient of each sound source in the next time period, and enhance the sound source.

[0129] In this example, the generalized sidelobe canceller GSC uses a fixed beamformer and adopts the minimum variance distortionless response MVDR method to generate a fixed beam. The expression of the beam output is:

[0130] (20)

[0131] Among them, represents the conjugate transpose of the head-related function of the q-th sound source, represents the inverse of the noise correlation matrix, represents the input microphone signal, and the denominator plays a role in normalization and weighting. It adjusts the input signal according to the sound source direction and noise characteristics, so that when the beam is output, the signal in the target sound source direction can be enhanced, and the signals in non-target directions (including noise) are suppressed. Compared with the traditional delay-and-sum beam, this method considering the interaural time difference is more refined. In binaural hearing, due to the existence of the head, there is a time difference in the arrival of sound at the two ears. Utilizing this characteristic can more accurately locate the sound source and improve the noise suppression effect.

[0132] The adaptive blocking matrix coefficient of each sound source in the next time period is the output result obtained by using the microphone signal in the next time period and the dynamic existence probabilities of each sound source as the input of the generalized sidelobe canceller. The specific acquisition principle of the adaptive blocking matrix coefficient is as follows:

[0133] In this embodiment, the generalized sidelobe canceller GSC uses an adaptive blocking matrix to suppress the sound in the target direction and generate a noise and interference reference signal for noise cancellation. For the target sound source q, first, a target signal projection matrix is constructed using the target existence probability result estimated by the Gaussian mixture model.

[0134] (21)

[0135] Among them, is the probability of the existence of the sound source q under the condition of the sound source direction The probability of the existence of the sound source q under the condition of the sound source direction, is the probability of the existence of the sound source calculated based on the signal-to-noise ratio. In a complex acoustic environment, the presence, intensity, and location of the sound source are dynamically changing. Through the calculation of the speech separation SPP algorithm and the Gaussian mixture model GMM, the probability of the existence of each sound source at different time-frequency frames can be estimated more accurately. This probability value plays a key weighting role in the construction of the target signal projection matrix. When is relatively high, it indicates that the possibility of the existence of the sound source in the current frame is relatively large. At this time, when constructing the projection matrix, more attention will be paid to using the information of the current frame to highlight the target sound source; conversely, when the probability is relatively low, more reference will be made to the projection matrix of the previous frame to maintain a certain stability and continuity.

[0136] represents the microphone signal covariance matrix, which reflects the correlation and energy distribution of the signals received by the microphones among different channels. represents the energy of the input signal. By taking as a part to participate in the calculation, the covariance matrix can be normalized, so that signals with different energy levels can be reasonably processed in the construction of the projection matrix, avoiding processing biases caused by too large differences in signal energy.

[0137] Calculate the complementary space matrix of the target signal, which shields the target signal.

[0138] (22)

[0139] For the sound source q, select the matrix parameters in the (M - 1)-th row and M-th column of the complementary space matrix to generate the adaptive blocking matrix coefficient. represents an identity matrix, the elements on its diagonal are 1, and the rest are 0. By subtracting the target signal projection matrix from the identity matrix, the obtained target signal complementary space matrix can shield the target signal. In subsequent processing, such as calculating the blocking matrix coefficient, the complementary space matrix plays a key role. It can extract the part orthogonal to the target signal, so as to better generate the noise and interference reference signals. For example, in an environment with multiple sound sources and noise, the target signal complementary space matrix can exclude the space where the target sound source signal is located, so that subsequent processing can focus more on the analysis and processing of noise and other interference signals, providing a more accurate basis for noise cancellation.

[0140] In this embodiment, the generalized sidelobe canceller (GSC) uses an adaptive noise canceller. Taking the output of the above-mentioned adaptive blocking matrix as a reference and the fixed beam output as the input signal, it further eliminates the noise and interference signals in the target direction. The adaptive noise canceller adopts a normalized least mean square filter, and the update process of the filter coefficients is as follows:

[0141] Taking the microphone signal in the current period as the input of the minimum variance distortionless response (MVDR) algorithm, the beam in the current period is output.

[0142] Among them, the MVDR algorithm is a well-known technology, and its specific principle will not be elaborated here.

[0143] The adaptive noise cancellation coefficient of the sound source q in the next period has the following expression:

[0144] (23)

[0145] In the formula, represents the adaptive noise cancellation coefficient of the sound source q in the current period, where its initial value is 0; represents the beam in the current period; represents the adaptive blocking matrix coefficient of the sound source q in the next period; represents the microphone signal in the current period; represents taking the absolute value.

[0146] Among them, the expression of the iteration step size of each sound source at the sound source angle in the current period is:

[0147] (24)

[0148] In the formula, represents the iteration step size of the sound source q at the sound source angle in the current period; represents the existence probability of the sound source q at the sound source angle in the current period; represents a preset value.

[0149] It should be noted that the value range of is generally between 0 and 1. In this embodiment, the value of

[0150] Further, the microphone signal in the next time period is used as the input of the minimum distortion variance, and the preliminarily enhanced signal is output. Then, the enhanced signal is used as the input of the generalized sidelobe canceller. Among them, the adaptive noise cancellation coefficients of each sound source in the next time period are used as the adaptive filtering coefficients in the adaptive filter of the generalized sidelobe canceller, and the finally enhanced signal is output.

[0151] Among them, the working principle of the generalized sidelobe canceller is a well-known technology, and the specific principle will not be elaborated here.

[0152] It should be noted that: the above-mentioned sequence of the embodiments of the present application is only for description and does not represent the advantages and disadvantages of the embodiments. And the above-mentioned specific embodiments of this specification have been described. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0153] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.

[0154] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; modifying the technical solutions recorded in the foregoing embodiments, or equivalently replacing some of the technical features, does not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A sound source localization and enhancement method for binaural hearing aids, characterized in that, The method includes the following steps: There are multiple sound sources in the environment. Obtain the received signals of each microphone within a preset duration, and divide the preset duration into multiple time segments; According to the correlation between the received signals of different microphones in any time segment, and combined with the random maximum likelihood estimation, determine the sound source angle in any time segment; At the sound source angle of any time segment, by analyzing the distribution of each sound source, and combined with the Gaussian mixture model, determine the existence probability of each sound source at the sound source angle of any time segment; Expression for the probability of the presence of sound source q at the sound source angle in time period b is as follows: ; where, represents the prior probability of the q-th Gaussian component in the Gaussian mixture model, where one Gaussian component represents one sound source; represents the q-th Gaussian component; represents the sound source angle in time period b; and represent the means of the q-th Gaussian component and the i-th Gaussian component in the Gaussian mixture model, respectively; and represent the variances of the q-th Gaussian component and the i-th Gaussian component in the Gaussian mixture model, respectively; represents the Gaussian distribution function; Q represents the number of all Gaussian components; Based on the existence probability of each sound source at the sound source angle of any time segment, determine the dynamic existence probability of each sound source at the sound source angle of any time segment; Based on the received signals of each microphone in the current time segment, obtain the beam of the current time segment; Based on the dynamic existence probability of each sound source at the sound source angle of the current time segment, determine the iteration step of each sound source at the sound source angle of the current time segment, and combined with the beam of the current time segment and the received signals of each microphone, to determine the adaptive noise cancellation coefficient of each sound source in the next time segment, and enhance the sound source; The method for determining the dynamic existence probability of each sound source is: Take the existence probability of each sound source at the sound source angle of any time segment as the input of the speech separation multi-channel speech presence probability SPP algorithm, and output the dynamic existence probability of each sound source at the sound source angle of any time segment; The expression for the iteration step size of each sound source at the sound source angle in the current time period is as follows: ; In the formula, represents the iteration step size of sound source q at the sound source angle in the current time period; represents the existence probability of sound source q at the sound source angle in the current time period; represents a preset value.

2. The sound source localization and enhancement method for a binaural hearing aid according to claim 1, characterized in that, The method for determining the sound source angle in any time segment is: Take the set composed of the received signals of all microphones in any time segment as the microphone signal, obtain the covariance matrix of the microphone signal, as the input of the random maximum likelihood estimation algorithm, iterate the random maximum likelihood estimation algorithm, and output the sound source angle in any time segment.

3. A sound source localization and enhancement method for a binaural hearing aid according to claim 1, characterized in that, The prior probability, mean and variance of the Gaussian components are all estimated by the expectation maximization algorithm.

4. The method for sound source localization and enhancement for a binaural hearing aid according to claim 2, characterized in that, The process of obtaining the beam of the current time segment is: Take the microphone signal in the current time segment as the input of the minimum variance distortionless response algorithm, and output the beam in the current time segment.

5. The method for sound source localization and enhancement for a binaural hearing aid according to claim 4, characterized in that, Determining the adaptive noise cancellation coefficient of each sound source in the next time segment includes: The adaptive noise cancellation coefficient of sound source q in the next time period has the following expression: ; where represents the adaptive noise cancellation coefficient of sound source q in the current time period, where has an initial value of 0; represents the beam in the current time period; represents the adaptive blocking matrix coefficient of sound source q in the next time period; represents the microphone signal in the current time period; represents taking the absolute value.

6. The sound source localization and enhancement method for a binaural hearing aid according to claim 5, characterized in that The adaptive blocking matrix coefficient of sound source q in the next time segment is the output result obtained by taking the microphone signal in the next time segment and the dynamic existence probability of sound source q as the input of the generalized sidelobe canceller.

7. The sound source localization and enhancement method for a binaural hearing aid according to claim 1, characterized in that, Enhancing the sound source includes: taking the microphone signal in the next time segment as the input of the minimum distortionless variance, outputting the preliminarily enhanced signal, and taking the enhanced signal as the input of the generalized sidelobe cancellation technique, where the adaptive noise cancellation coefficient of each sound source in the next time segment is used as the adaptive filtering coefficient in the adaptive filter of the generalized sidelobe canceller, and output the finally enhanced signal.

Citation Information

Patent Citations

  • A microphone system and a hearing device comprising same

    CN109040932A

  • De-noising method for multi-microphone audio equipment, in particular for a "hands-free" telephony system

    US20120322511A1