Robust binaural beam forming method
By adopting a robust binaural beamforming method in hearing aids, the covariance matrix is calculated and spatial filter is designed, the spatial clue retention problem caused by HRTF mismatch is solved, and binaural clue retention and noise reduction balance in the case of HRTF mismatch is achieved.
Patent Information
- Application Number
- CN202510245767.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-03
AI Technical Summary
Traditional hearing aids are difficult to effectively retain spatial clues in the case of HRTF mismatch, resulting in the wearer being unable to accurately sense the direction and position of the sound source.
Using a robust binaural beamforming method, by calculating the covariance matrix of the noise-containing signals received by all microphones in the binaural hearing aid, and using the Cauchy–Schwarz inequality and singular value decomposition method, a spatial filter is designed to minimize the binaural cues loss of the desired voice source, retain binaural cues as much as possible, and the balance of noise reduction and binaural cues retention is achieved through parameter α.
In the case of HRTF mismatch, binaural cues of the desired voice source can be retained as much as possible, and the balance between noise reduction and binaural cues retention can be achieved, and the wearer's perception of the direction and position of the sound source can be improved.
Smart Images

Figure CN120091254A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of acoustic signal processing, and more particularly, to a robust binaural beamforming method. Background Art
[0002] Nowadays, digital hearing aids have attracted more and more attention. As an important auxiliary tool, it can play a positive role in improving the hearing conditions of hearing-impaired people. Among them, the binaural beamforming acoustic technology plays an important role in the field of hearing aids and becomes one of the key technologies for hearing aids to improve the hearing effect. On the one hand, the binaural beamforming technology can suppress noise and improve the clarity of speech. On the other hand, it can retain the spatial cues of sound, and the wearer can use these spatial cues to perceive the direction and specific location of the sound source.
[0003] The Head-Related Transfer Function (HRTF) in binaural beamforming plays a key role in the spatial perception of sound. The HRTF is the acoustic transfer function of sound propagating from different directions to the ears, which helps the wearer retain the spatial cues of sound and thus helps the human ear perceive the spatial position of sound. However, whether it is the current estimation of HRTF or the use of the HRTF dataset, it is inevitable to introduce a certain degree of error. When estimating the HRTF, it is difficult to be completely accurate due to many factors, such as the limitations of measurement equipment. When using the HRTF dataset, due to the differences in physiological structures and other aspects among different individuals, the HRTF in the dataset is often not exactly the same as the actual wearer's own true HRTF, and thus corresponding errors will occur. These errors will cause deviations when the binaural beamforming retains spatial cues, resulting in the wearer being unable to accurately perceive the direction and position of the sound source. Summary of the Invention
[0004] In view of this, this application provides a robust binaural beamforming method to solve the problem that traditional hearing aids are difficult to effectively retain spatial cues when there is an HRTF mismatch problem.
[0005] To achieve the above object, the technical solution adopted in this application is as follows: A robust binaural beamforming method, comprising: Step 1: Calculate the noisy signals received by all microphones in the binaural hearing aids (HA) according to formula (2); (2) Wherein, y= Y 1 , Y 2 , …,Y M T represents the noisy signals received by all microphones, [ ] T represents the transpose of a vector or matrix, S represents the expected speech source signal after Fourier transform, B= b 1 , b 2 , …,b P represents the HRTF combination of all interference sources, b j = b j1 , b j2 , …,b jM T represents the j th HRTF of the interference source, u= U 1 , U 2 , …,U P T represents P interference source signals, x = X 1 , X 2 , …,X M T represents the expected speech source signals received by all microphones, z j = Z j1 , Z j2 , …, Z jM T represents the j th interference source signal received by all microphones, v = V 1 , V 2 , …,V M T Represents the background noise of all microphones, represents the HRTF of the desired speech source for all microphones, represents a complex number, M represents the total number of microphones equipped in the binaural hearing aid, where M is an even number and each ear is equipped with M / 2 microphones; Step 2: Assume that the desired speech source, interference source, and noise are uncorrelated, and calculate the covariance matrix of the noisy signal y and the covariance matrix of the total noise R yy according to formulas (3) and (4). The total noise is the noise plus the interference source signal; R nn (3) R nn = R zz + R vv (4) where, {} is the expectation operator, and [ ] H represents the conjugate transpose of a vector or matrix, R xx represents the covariance matrix of the desired speech source signal, represents the covariance matrix of the interference source signal, represents the noise covariance matrix. Calculate R xx according to the Target Activity Detector (TAD), and since the desired speech source, interference source, and noise are uncorrelated, obtain the covariance matrix of the total noise (noise plus interference source signal) R nn = R zz + R vv . The signal-to-interference-plus-noise ratio of the input signal (dB), where, and are respectively M -dimensional vectors composed of elements of 0 and 1, that is and ; represents the first element in the vector, The M th element in the vector;
[0006] Step 3: Calculate the binaural cue loss of the desired speech source and the j th interfering source according to formula (8): (8) where a L and a R represent the HRTFs of the desired speech source for the left and right ear reference microphones respectively, and represent the binaural cues before and after filtering by the spatial filter respectively; and represent the spatial filters at the left ear and the right ear respectively; [[ ]] H represents the conjugate transpose of a matrix or vector; Step 4: If the HRTF of the desired speech source obtained through experiments and calculations is , in the case of HRTF mismatch, the obtained HRTF of the desired speech source is not completely accurate. Assuming the HRTF error vector of the desired speech source is e , then the exact HRTF of the desired speech source is , e is an M-dimensional vector. Calculate the worst-case binaural cue retention of the desired speech source according to formula (13): (13) where P a is the worst-case binaural cue retention of the desired speech source, e L and e R are the HRTF errors of the desired speech source for the left and right ear hearing aid reference microphones respectively.
[0007] Step 5: By making the e two-norm less than or equal to a constant η , that is , and using the Cauchy–Schwarz inequality method, further simplify formula (13) in Step 4 to get: (14) Step 6: According to the total noise covariance matrix Rnn and the P a , when the HRTF is mismatched, minimize the loss of binaural cues of the desired speech source through the filter w in the worst case of binaural cue retention, so as to retain the binaural cues of the desired speech source as much as possible. According to different individual adaptation situations, balance noise reduction and binaural cue retention by selecting the parameter α, and list the cost function: (15) where , , is a zero matrix of dimension
[0008] Step 7: Decompose the matrix by the singular value decomposition method, , and reorganize the optimization problem: (16) where Tr ( ) represents the rank of the matrix. The CVX toolbox is a tool in Matlab that can solve convex optimization problems. Since the above problem is a convex optimization problem, the CVX toolbox in Matlab can be used to solve the filters w L and w R .
[0009] Compared with the prior art, the beneficial effects of this application are: 1. The method of this application can retain the binaural cues of the desired speech source as much as possible by using the second norm and the Cauchy–Schwarz inequality in the case of HRTF mismatch, and achieve a balance between noise reduction and binaural cue retention by adopting parameters.
[0010] 2. For the HRTF mismatch problem, the method of this application can be applied to process signals in hearing aids. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of this application, so they should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 is a flowchart of a robust binaural beamforming method in this application; Figure 2 is a schematic diagram of a binaural hearing aid configuration in this application; Figure 3 Schematic diagram of the experimental configuration of the binaural hearing aid in this application; Figure 4 Comparison graph of the output SINR of the robust binaural beamforming method and BMVDR and PUB in this application; Figure 5 Schematic diagram of the retention of binaural cues of the desired speech source when η = 0.005 in this application; Figure 6 Schematic diagram of the retention of binaural cues of the desired speech source when α = 100 in this application; Figure 7 Schematic diagram of the output SINR of the binaural filter when α = 100 in this application. Detailed implementation manners
[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application.
[0014] As Figure 1 and Figure 2 shown, Step 1: Calculate the noisy signals received by all microphones in the binaural hearing aids (HA); In this application, it is assumed that a total of M microphones are equipped for the binaural hearing aids (HA) (where M is an even number and each of the left and right ears is equipped with M / 2 microphones, as Figure 2 shown). Considering a scenario with a single desired speech source and P non-coherent interference sources, in the short-time Fourier transform (STFT) domain, let l be the frame index, ω be the angular frequency, then the noisy signal received by the k th microphone is:
[0015] (1) where, a k ( ω ) represents the HRTF of the desired speech signal relative to the k th microphone, S ( ω , l ) represents the desired speech source, b jk (ω ) is the HRTF from the j th interfering source to the k th microphone, U j ( ω , l ) is the j th interfering source, V k ( ω , l ) represents the background noise of the k th microphone. X k ( ω , l ) = a k ( ω ) S ( ω , l ), Z jk ( ω , l ) = b jk ( ω ) U j ( ω , l ), X k ( ω , l ) represents the value of the desired speech source signal received by the k th microphone in the Short-Term Fourier Transform (STFT) domain, Z jk ( ω , l ) is the value of the k th interfering source signal received by the j th microphone in the STFT domain.
[0016] Since all beamforming operations are performed frame-by-frame and frequency-by-frequency in the STFT (Short-Time Fourier Transform) domain, for the sake of concise writing, the frequency index ω and the frame index l can be omitted. Accordingly, Equation (1) can be written compactly as:
[0017] (2) where, y= Y 1 , Y 2 , …,Y M T represents the signals received by all microphones, T represents the transpose of a vector or matrix, S represents the desired speech source signal after Fourier transform, B= b 1 , b 2 , …,b P represents the HRTF combination of all interfering sources, b j = b j1 , b j2 , …,b jM T represents the j th HRTF of the interfering source, u= U 1 , U 2 , …,U P T represents P interfering source signals, x = X 1 , X 2 , …,X M T represents the desired speech source signals received by all microphones, z j = Z j1 , Z j2 , …, Z jM T represents the j th interfering source signals received by all microphones, v = V 1 , V 2 , …,V M T Represents the background noise of all microphones, represents the HRTF of the desired speech source for all microphones, represents a complex number.
[0018] Step 2: Assume that the desired speech source, interfering sources, and noise are uncorrelated, and calculate the covariance matrix of the noisy signal y according to formulas (3) and (4). R yy and the covariance matrix of the total noise R nn , where the total noise is the noise plus the interfering source signals; (3) R nn = R zz + R vv (4) where {} is the expectation operator, and [ ] H represents the conjugate transpose of a vector or matrix, R xx represents the covariance matrix of the desired speech source signal, represents the covariance matrix of the interfering source signals, represents the noise covariance matrix. Calculate R xx using the Target Activity Detector (TAD), and since the desired speech source, interfering sources, and noise are uncorrelated, obtain the covariance matrix of the total noise (noise plus interfering source signals) R nn = R zz + R vv . The signal-to-interference-plus-noise ratio of the input signal (dB), where and are M-dimensional vectors consisting of 0 and a single element of 1, i.e., and , represents the first element in the vector, represents M the
[0019] Specifically, the Rxx , R zz and R vv are calculated by the following equations: (5) (6) (7) where P s and P uj are the variances of the desired speech source and the j th interfering source, respectively, and represents the HRTF of the j th interfering source.
[0020] Step 3: Calculate the binaural cue loss of the desired speech source and the j th interfering source according to Equation (8): (8) where a L and a R represent the HRTFs of the desired speech source for the left and right ear reference microphones, respectively, and represent the binaural cues before and after filtering by the spatial filter, respectively; and represent the spatial filters at the left ear and the right ear, respectively; H represents the conjugate transpose of a matrix or vector.
[0021] To perform noise reduction and binaural cue preservation, two spatial filters need to be designed, namely and . In this application, it is assumed that the noisy signals received by the microphones at each HA can be transmitted to the other HA through a lossless communication channel. Each hearing aid has an M-dimensional signal y . For a hearing-impaired listener, the designed hearing aid filter not only needs to suppress environmental noise but also needs to preserve the spatial cues of the sound source. To achieve noise reduction and binaural cue preservation, two spatial filters need to be designed, denoted as and respectively, and applied to the reference microphones. Let the first microphone and the last microphone be the reference microphones on the left and right ear hearing aids, respectively. For clarity of expression, the subsequent subscripts L and R will be used to represent these reference microphones. Applying the binaural beamformer, the output signals of the left and right ear hearing aids are given by the following equations:
[0022] (9) Generally, the human ear judges the spatial position of a sound source based on the spatial cues of binaural audio signals. A classic method for preserving binaural cues is through the Interaural Transfer Function (ITF), which is defined as the ratio of the HRTFs of the reference microphones for the two ears, i.e.:
[0023] (10) where a L and a R represent the HRTFs of the desired speech source for the left and right ear reference microphones respectively, and represents the input ITF of the desired speech source.
[0024] Applying a binaural beamformer, the filtered binaural cues are given by the following equation: (11) From the above equations (10) and (11), the binaural cue loss of the desired speech source and the j th interfering source is calculated as: (12) If is equal to 0, it means that the binaural cues of the desired speech source and the j th interfering source are completely preserved.
[0025] The following steps are the method for solving the filters and .
[0026] If , it means that the binaural cues of the desired speech source are completely preserved. In the case of HRTF mismatch, it is difficult to completely preserve the binaural cues. Therefore, it is necessary to preserve the binaural cues of the desired speech source as much as possible.
[0027] Step 4: If the HRTF of the desired speech source obtained through experiments and calculations is , in the case of HRTF mismatch, the obtained HRTF of the desired speech source is not completely accurate. Assuming that the HRTF error vector of the desired speech source is e , then the accurate HRTF of the desired speech source is , e is an M-dimensional vector. Calculate the worst-case scenario of binaural cue preservation for the desired speech source according to Equation (13): (13) where P a is the worst-case scenario of binaural cue preservation for the desired speech source,e L and e R are the HRTF errors of the desired speech source with respect to the reference microphones of the left and right ear hearing aids, respectively.
[0028] Step 5: By making e the two-norm less than or equal to a constant η , that is , and using the Cauchy–Schwarz inequality method, the formula (13) in Step 4 is further simplified to obtain: (14) Step 6: According to the total noise covariance matrix R nn obtained in Step 2 and P a obtained in Step 5, the cost function is listed as shown in Equation (15); (15) The said cost function is used to calculate the value of the filter w when the value is minimized. The parameter is used to adjust the balance between noise reduction and binaural cue retention. represents the noise power after filtering. By filtering, the noise power is minimized, that is, noise reduction. represents the degree of binaural cue retention dominated by the parameter . If is larger, then the cost function is more used to make minimal. If is extremely small, it means that the filter and the binaural cues after filtering are basically the same, that is, the binaural cues are basically retained. Among them, , is a zero matrix of dimension w . When the HRTF is mismatched, the loss of the binaural cues of the desired speech source is minimized through the filter α in the worst case of binaural cue retention, so as to retain the binaural cues of the desired speech source as much as possible. According to different individual adaptation situations, by selecting the parameter α , the balance between noise reduction and binaural cue retention is achieved. α The larger
[0029] Step 7: Decompose the matrix by the singular value decomposition method, , and reorganize and optimize the problem: (16) where Tr ( ) represents the rank of the matrix. The CVX toolbox is a tool in Matlab that can solve convex optimization problems. Since the above problem is a convex optimization problem, the CVX toolbox in Matlab can be used to solve the filters w L and w R .
[0030] Example: In this example, a total of M = 4 microphones are placed in a room with a specification of (10 × 10) m, as Figure 3 shown. By placing two microphones on each ear, the first microphone and the last microphone are used as the reference microphones for the two hearing aids. Assume that the user wearing the hearing aids is located at the center of the room, the radius of the user's head is 10 cm, and the reverberation time T60 = 200 ms. The desired speech source is set directly in front of the head (0°), and the desired speech source signal is a speech signal from the TIMIT database. The 4 interference sources are placed at 30°, 60°, -30°, -60° respectively, and they are also speech signals from the TIMIT dataset. All the signal sources are distributed on a circle centered on the head (radius = 1 m), the elevation angle is set to 0°, and they remain stationary. The input signal-to-noise ratio of the interference sources is set to 5.8 Hz, and the microphone self-noise is simulated with Gaussian white noise at 30 Hz. All signals are sampled at 16 kHz, and framed with 25% overlap using a 16-ms square root Hann window.
[0031] The simulation of the proposed method is compared with the reference methods BMVDR (Binaural Minimum Variance Distortionless Response) and PUB (Parametric Unconstrained Beamformer) in terms of binaural cue preservation and noise reduction performance. The noise reduction performance and binaural cue retention ability are compared by evaluating the binaural signal-to-interference plus noise ratio (SINR) and binaural cue loss. Among them, is the calculation method for the output SINR, It is a calculation method for the binaural cue loss of the desired speech source. This experiment only considers the mismatch of the HRTF of the desired speech source and the non-mismatch of the HRTF of the interfering source, and compares the retention of the binaural cues of the desired speech source and the overall SINR. Taking the desired speech source η as 0.001, as α increases, the output SINR of this method is compared with that of BMVDR and PUB, as Figure 4 shown. When η = 0.005, the retention of the binaural cues of the desired speech source is as Figure 5 shown. Figure 6 is the retention of the binaural cues of the desired speech source when α is 100 and there are different HRTF mismatches, where η increases as the degree of HRTF mismatch increases. Figure 7 is the output SINR when α = 100 and η increases as the degree of HRTF mismatch increases. Among them, prop in the figure is the method of this application, and unproc is the input SINR. The degree of signal distortion will increase as the degree of HRTF mismatch increases.
[0032] This application aims to retain spatial cues as much as possible in the case of HRTF mismatch, and at the same time achieve a balance between spatial cue retention and noise reduction effect through parameter adjustment. For the hearing characteristics of different individuals, flexibly select appropriate parameter settings to optimize the user's listening experience. This method shows unique advantages when compared with the BMVDR and PUB methods.
[0033] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should be covered by the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.
Claims
1. A robust binaural beamforming method, characterized in that: include: Step 1: Calculate the noisy signals received by all microphones in the binaural hearing aids according to formula (2); (2) in, y = [ Y 1, Y 2, …,Y M ] T Indicates the noisy signal received by all microphones, [ ] T represents the transpose of a vector or matrix, S represents the desired speech source signal after Fourier transform, B = [ b 1, b 2, …,b P ] represents the HRTF combination of all interference sources, b j = [ b j1 , b j2 , …,b jM ] T Indicates j The HRTF of the interference source, u = [ U 1, U 2, …,U P ] T express P Interference source signal, x = [ X 1, X 2, …,X M ] T represents the expected speech source signal received by all microphones, z j = [ Z j1 , Z j2 , …, Z jM ] T Indicates the first j Interference source signal, v = [ V 1, V 2, …,V M ] T represents the background noise of all microphones, represents the HRTF of the expected speech source for all microphones, M Indicates the total number of microphones equipped in binaural hearing aids, where M An even number of 1 for each ear M / 2 microphones, Indicates plural; Step 2: Assuming that the desired speech source, interference source, and noise are uncorrelated, calculate the noisy signal according to formulas (3) and (4): y The covariance matrix of R yy and the covariance matrix of the total noise R nn , the total noise is the noise plus the interference source signal; (3) R nn = R zz + R vv (4) in, {} is the expectation operator, [ ] H represents the conjugate transpose of a vector or matrix, R xx represents the covariance matrix of the desired sound source signal, represents the covariance matrix of the interference source signal, Represents the noise covariance matrix. Calculated according to the target activity detector R xx , and since the expected speech source, interference source and noise are uncorrelated, the total noise covariance matrix is obtained R nn = R zz + R vv ; Signal-to-interference-noise ratio of the input signal (dB), where and is a vector consisting of 0 and one 1, that is and , express The first element in the vector, express The first M elements; Step 3: Calculate the expected speech source and the j Loss of binaural cues for distractors: (8) in, a L and a R They are respectively represented as the HRTF of the desired speech source for the left and right ear reference microphones, and Respectively represent binaural cues before and after filtering by spatial filter; and denote the spatial filters at the left ear and the right ear respectively; [ ] H Represents the conjugate transpose of a matrix or vector; Step 4: In the case of HRTF mismatch, assume that the expected speech source and the j The precise HRTF of the interference source is , the estimated HRTF is , e is the expected speech source HRTF error vector, e is an M-dimensional vector, and the worst case of binaural clue preservation of the expected speech source is calculated according to formula (13): (13) in, P a Keep the worst case binaural cues for the desired speech source, e L and e R are the HRTF errors of the reference microphones of the left and right hearing aids for the desired speech source, respectively; Step 5: By e The bi-norm of is less than or equal to a constant η ,Right now , and using the Cauchy–Schwarz inequality method, further simplify formula (13) in step 4 to obtain: (14) ; Step 6: The total noise covariance matrix obtained in step 2 R nn and the value obtained in step 5 P a , list the cost function; (15) in, , , for dimensional all-zero matrix; Step 7: Use the singular value decomposition method to transform the matrix To decompose, , re-arrange the optimization problem and use the CVX toolbox in Matlab to solve the filter w L and w R ; (16) in, Tr ( ) represents the rank of the matrix.