Estimating direct-to-reverberation ratio of a sound signal
By using a single-microphone hearing device approach, combined with discrete Fourier transform and machine learning algorithms, the high computational cost and reliance on sound field assumptions in existing technologies are solved, achieving low-cost and efficient direct reverberation ratio estimation, applicable to complex sound fields.
Patent Information
- Application Number
- CN202110148911.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-06
- Filing Date
- 2021-02-03
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2041-02-03
AI Technical Summary
Existing hearing devices suffer from problems such as high computational cost, strong dependence on sound field assumptions, and the need for multiple microphones when estimating direct reverberation ratio, making it difficult to accurately estimate direct reverberation ratio in complex sound fields.
A single-microphone-based approach is used to determine the acoustic start signal by performing discrete Fourier transform and time frame analysis on the sound signal, and to estimate the direct reverberation ratio using machine learning algorithms such as linear regression models or artificial neural networks.
It achieves direct reverberation ratio estimation with low computational cost, suitable for online and offline applications, independent of signal level and microphone directivity, applicable to monoaural or binaural hearing devices, and improves the accuracy and flexibility of estimation.
Smart Images

Figure CN113299316B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to methods, computer programs, and computer-readable media for estimating the direct-to-reverberant ratio of sound signals. Furthermore, this invention relates to a hearing device. Background Technology
[0002] Hearing devices are typically small and complex. They may include a processor, microphone, speaker, memory, housing, and other electronic and mechanical components. Some examples of hearing devices are behind-the-ear (BTE) devices, in-the-canal (RIC) devices, in-the-ear (ITE) devices, completely in-the-canal (CIC) devices, and in-the-canal (IIC) devices. Users can choose one of these hearing devices based on hearing loss, aesthetic preferences, lifestyle needs, and budget, comparing it to another device.
[0003] Everyday sounds captured by hearing devices are constantly affected by reverberation. For users of hearing devices, reflected sound waves aid in spatial and distance perception. For algorithms that process sound waves using hearing devices, knowledge of the amount of reverberation present in the sound signal generated from the sound waves can be beneficial.
[0004] Several methods have been proposed for estimating the direct reverberation (energy) ratio (DRR). However, these methods may have drawbacks when used in hearing device applications. Furthermore, all methods are based on assumptions about the sound field, which are not always satisfied in reality. Some methods rely on the assumption of an isotropic sound field. Some methods require prior knowledge of the direction of arrival of the sound source. In all these cases, at least one microphone is used.
[0005] US 20170303053 A1 relates to a hearing device in which a déroresistive process is performed, wherein the déroresistive process measures a dedicated reverberation reference signal to determine the reverberation characteristics of an acoustic environment, and reduces the reverberation effect in the output signal of the hearing device based on the reverberation characteristics. Summary of the Invention
[0006] The object of this invention is to provide a method for estimating direct reverberation ratio, suitable for hearing device applications. Another object of this invention is to provide a method for estimating direct reverberation ratio that has low computational cost, is easy to implement, and can be performed using sound signals recorded by only one microphone.
[0007] These objectives are achieved through the subject matter of the independent claims. Other exemplary embodiments will be apparent from the dependent claims and the following description.
[0008] A first aspect of the invention relates to a method for estimating the direct reverberation ratio of an acoustic signal. This method can be performed by a hearing device. The hearing device may include a microphone that generates the acoustic signal. The hearing device may be worn by a user, for example, behind or in the ear. The hearing device may be a hearing aid for compensating for a user's hearing loss. Herein and below, when referring to a hearing device, it also means a pair of hearing devices, i.e., hearing devices for each of the user's ears. The hearing device may include hearing aids and / or cochlear implants.
[0009] The direct reverberation ratio, or more precisely, the direct reverberation energy ratio, is the ratio of the direct sound received from the sound source to the reverberated sound received from reflections in the environment of the sound source.
[0010] Direct sound can be based on sound waves that travel directly from one or more sound sources to a microphone that acquires the sound signal. Reflected and / or reverberated sound can be sound waves from one or more sound sources that are reflected in the environment. The direct reverberation ratio can be a number, for example, between 0 and 1, where 0 can represent sound without reverberation, and / or 1 can represent sound with only reverberation. The direct reverberation ratio can also be provided in dB.
[0011] According to an embodiment of the present invention, the method includes: determining a first energy value of an audio signal in a first time frame. The audio signal can be determined within the time frame. For each time frame, at least one energy value can be calculated based on the audio signal. The time frames may all have equal length. The time frames may overlap. The energy value may indicate the energy of the audio signal or at least one frequency band of the audio signal in the corresponding time frame.
[0012] For example, a discrete Fourier transform can be performed on an audio signal. Specifically, the audio signal can be time-buffered, overlapped, windowed, and then subjected to a Fourier transform. Power estimation for each frame can then be performed. The audio signal can be divided into time frames, and within each time frame, the audio signal is transformed into frequency bins, which indicate the intensity of the audio signal within the frequency range associated with that bin. Based on these intensities (i.e., the Fourier coefficients), the energy value can be calculated.
[0013] According to an embodiment of the invention, the method further includes: assigning a positive value to the onset value of the first time frame if the difference between the first energy value of the first time frame and the second energy value of the previous second time frame is greater than a threshold; otherwise, assigning a zero value. At least one onset signal can be determined based on one or more energy values. When the energy value of a time frame is higher than the energy value of the previous time frame by more than a threshold, the onset value for the onset signal of the time frame can be set to a positive value. Otherwise, the onset value is set to zero. A start, or more specifically an acoustic start, can be defined as a sudden jump in the energy of a sound signal, particularly an upward jump.
[0014] The start signal can include a start value for each time frame. A positive value can indicate the presence and / or magnitude of a start. To determine the start and / or start value, the energy value of the time frame is compared with the energy value of a previous time frame. A start is assumed to exist when the difference between the energy value of the time frame and the energy value of a previous time frame exceeds a threshold.
[0015] When a start is detected for a time frame, the start value is set to a positive value, which can be 1. Typically, the positive value can be higher than the threshold value, which is zero.
[0016] When a start is detected for a time frame, the start value is set to a positive value, or it can be the difference between the energy value in the time frame and the energy value in the previous time frame. When no start is detected, the start value can be set to 0.
[0017] Typically, this method is based on the effect of reverberation on acoustic onset. Reverberation can generally contaminate the spectrum of a sound signal. Therefore, it can be assumed that the number and intensity of acoustic onsets decrease with increasing reverberation.
[0018] It is important to note that more than one energy value can be determined for each time frame for different properties of the sound signal (e.g., different frequency bands). Then, more than one start signal can be determined for each property individually.
[0019] According to an embodiment of the invention, the method further includes: determining the direct reverberation ratio by providing a start signal, including a start value, to a machine learning algorithm, the machine learning algorithm having been trained to determine the direct reverberation ratio based on the start signal. The direct reverberation ratio can be determined by inputting at least one start signal and / or features derived therefrom into the machine learning algorithm, the machine learning algorithm having been trained to generate the direct reverberation ratio according to at least one start signal.
[0020] One or more start signals can be input into a machine learning algorithm. The input audio signal can be preprocessed before being fed into the machine learning algorithm. For example, as described below, the start signal can be integrated and / or the gradient of the integrated start signal can be determined. The integrated start signal and / or gradient can then be input into the machine learning algorithm.
[0021] The machine learning algorithm has been trained to determine the direct reverberation ratio based on one or more start signals. Typically, the machine learning algorithm may have parameters (e.g., weights or coefficients) that have been adapted during training, such that when one or more start signals and / or parameters derived therefrom are input together with a known direct reverberation ratio, the machine learning algorithm outputs that direct reverberation ratio.
[0022] The method described in this paper for determining one or more start signals from a single sound signal and then using a machine learning algorithm to determine the direct reverberation ratio is easy to implement and requires relatively little computation, depending on the choice of a suitable machine learning algorithm. It should be noted that fairly simple machine learning algorithms, such as regression models, can be used.
[0023] By selecting an appropriate positive value for the start signal, this method can operate independently of the signal level, i.e., regardless of the recorded loudness. This method is applicable to both online and offline applications. It is efficient in terms of required memory and power. This method can be used monoaurally or binaurally.
[0024] This method does not require prior knowledge of the directional angle of the incoming sound. Furthermore, it is unaffected by the microphone's directional pattern.
[0025] According to embodiments of the present invention, machine learning algorithms are trained relative to the type of hearing device. Training data can be recorded and generated for specific types of hearing devices with specific hardware (e.g., housing and / or microphone and / or microphone position). Machine learning algorithms can also be trained differently for hearing devices in the left and right ears.
[0026] According to an embodiment of the invention, the start signal is integrated over time to determine the gradient of the start signal and provide this gradient to a machine learning algorithm. Integration and / or determination of the gradient for each start signal can be performed on one or more start signals. Integration can be performed over a time interval that begins at a specific time point and ends at the time point where the start value of the integration is determined. The gradient of each start signal can then be input into the machine learning algorithm. As already mentioned, one or more start signals can be preprocessed before being input into the machine learning algorithm.
[0027] The start signal can be integrated by summing the energy values of the time frames sorted by time. In other words, the value of the start signal as an integral over time frames can be the sum of the energy values of all previous time frames.
[0028] The gradient of the starting signal of the integral can be the average gradient of the starting signal of the integral. Such an average gradient can be determined based on the gradients of at least some points defined by the starting signal of the integral. It can also be determined by linear regression. Typically, the gradient can be a number indicating the rise of the corresponding starting signal.
[0029] According to embodiments of the present invention, a state-space model is used to determine the gradient of each starting signal. Using a state-space model, the gradient can be determined with less computational requirement because it may not be necessary to invert the matrix.
[0030] According to an embodiment of the invention, the machine learning algorithm is, or at least includes, a linear regression model. The direct reverberation ratio can then be determined based on the gradient of the starting signal of the integral. The gradient can be input into the linear regression model, which may include a linear function that weights the gradient and produces the direct reverberation ratio. The gradient weights may have been determined by training the machine learning algorithm.
[0031] It is important to note that other machine learning algorithms can also be used. For example, one or more start signals can be input into an artificial neural network that has been trained to classify start signals. The classifier output by the artificial neural network can be a direct reverberation ratio or a range of direct reverberation ratios.
[0032] There are several possible ways to utilize the properties of a sound signal to generate different energy values for each time frame. The total energy of the sound signal can be used. Alternatively, the energy of a frequency band of the sound signal can be used. As another possibility, loud and / or quiet sounds can be removed from the sound signal, and then the energy value can be determined based on the sound signal with the loud and / or quiet sounds removed.
[0033] According to embodiments of the present invention, a broadband energy value is determined for a first time frame or for each time frame, the broadband energy value indicating the energy of the sound signal within the time frame. For example, the broadband energy value can be determined based on all frequency boxes within the time frame. The energy value of a frequency box can be proportional to the square of the absolute value of a complex Fourier coefficient. These energy values can be summed.
[0034] According to an embodiment of the present invention, when the broadband energy value of a time frame is higher than the broadband energy value of a previous time frame by more than a broadband threshold, a broadband start signal is determined by setting the broadband start value of the broadband start signal for the time frame to a positive value. The broadband start signal can be determined based on the broadband energy value. As described above, the positive value can be set to 0 and 1. When the criteria for the start of a time frame are met, the positive value can also be set to the difference between the broadband energy value of that time frame and the broadband energy value of a previous time frame.
[0035] According to embodiments of the present invention, a frequency band energy value is determined for each time frame, the frequency band energy value indicating the energy of the sound signal in the frequency band of that time frame. Specific frequency boxes can be assembled into frequency bands, and the energy value of the frequency band can be determined solely based on the Fourier coefficients of the associated frequency boxes.
[0036] For example, a frequency band can have a lower limit that is above the mid-frequency range of the full spectrum available for a sound signal. Naturally, higher frequencies can be more affected than lower frequencies because reverberation-based sound diffraction primarily occurs in the high-frequency range.
[0037] According to an embodiment of the present invention, when the frequency band energy value of a time frame is higher than the frequency band energy value of a previous time frame by more than a frequency band threshold, the frequency band start signal is determined by setting the frequency band start value of the frequency band start signal for the time frame to a positive value.
[0038] The start signal of a frequency band can be determined based on the frequency band energy value. As described above, positive values can be set to 0 and 1. When the criteria for the start of a time frame are met, a positive value can also be set to the difference between the frequency band energy value of that time frame and the frequency band energy value of the previous time frame.
[0039] According to embodiments of the present invention, the sound signal is divided into multiple frequency bands, and a band start signal is determined for each frequency band. Frequency bands may overlap. Frequency bands may also cover the entire spectrum available for the sound signal.
[0040] According to embodiments of the present invention, the bandwidth threshold is different from the broadband threshold. For example, the bandwidth threshold is lower than the broadband threshold.
[0041] According to embodiments of the present invention, the frequency band thresholds are different for different frequency bands. For example, the frequency band threshold for lower frequencies is lower than the frequency band threshold for higher frequencies.
[0042] According to embodiments of the present invention, a broadband start signal and multiple band start signals are determined and input into a machine learning algorithm. These broadband start signals and multiple band start signals can cover the frequency range available from the sound signal. This can improve the accuracy of the direct reverberation ratio. The broadband start signal can be determined with a positive value set to 1, and multiple band start signals for multiple frequency bands can be determined with a positive value set to 1.
[0043] It is also possible to determine different band start signals for the same frequency band, which is determined in different ways (e.g., using different types of positive values). Multiple first band start signals for multiple frequency bands can be determined when the positive value is set to 1. Furthermore, multiple second band start signals for multiple frequency bands can be determined when the positive value is set to the difference between the energy value in a time frame and the energy value in a previous time frame.
[0044] The two previous embodiments can be combined, that is, the broadband start signal, the first frequency band start signal and the second frequency band start signal can be determined and input into a machine learning algorithm.
[0045] Another aspect of the invention relates to a method for operating a hearing device, the method comprising: generating a sound signal using a microphone of the hearing device; estimating the direct reverberation ratio of the sound signal as described above and below; processing the sound signal using the direct reverberation ratio to compensate for hearing loss in a user of the hearing device; and outputting the processed sound signal to the user. The direct reverberation ratio can be determined using a software module running in a processor of the hearing device. The sound signal processing can be performed using a sound processor of the hearing device, which can be tuned using the direct reverberation ratio.
[0046] According to embodiments of the invention, the direct reverberation ratio is used for at least one of the following: noise cancellation, reverberation cancellation, frequency-dependent amplification, frequency compression, beamforming, sound classification, self-voice detection, and foreground / background classification. Each of these functions can be performed using software modules (e.g., programs for hearing devices). These software modules can use the direct reverberation ratio as an input parameter.
[0047] For example, noise cancellation algorithms can better estimate background noise based on the direct reverberation ratio. Reverberation cancellation can benefit for the same reason. Gain models (i.e., frequency-dependent amplification) and / or compressors (i.e., frequency compression) can be better tuned based on the amounts of direct and reverberant energy, which can be determined from the direct reverberation ratio. Adaptive beamformers may have better noise reference estimates based on the direct reverberation ratio. Sound classifiers can also be improved by using the direct reverberation ratio as an additional input parameter. In particular, procedures for “speech in reverberation” can be optimized by additionally inputting the direct reverberation ratio.
[0048] Another aspect of the invention relates to a computer program for estimating the direct reverberation ratio of a sound signal and optionally for operating a hearing device, the computer program being adapted, when executed by a processor, to perform the steps of the methods described above and below. Another aspect of the invention also relates to a computer-readable medium in which such a computer program is stored.
[0049] For example, a computer program can be executed in the processor of a hearing device, which may be worn behind the ear by a person. The computer-readable medium may be the memory of the hearing device.
[0050] Typically, computer-readable media can be floppy disks, hard disks, USB (Universal Serial Bus) storage devices, RAM (Random Access Memory), ROM (Read-Only Memory), EPROM (Erasable Programmable Read-Only Memory), or flash memory. Computer-readable media can also be data communication networks that allow the download of program code, such as the Internet. Computer-readable media can be non-transitory or transient.
[0051] Another aspect of the invention relates to a hearing device adapted to perform the methods described above and below. The hearing device may include a microphone, a sound processor, a processor, and a sound output device. The method can be readily integrated into the hearing device because it can utilize features already available in the hearing device's DSP block and / or sound processor.
[0052] A microphone can be used to acquire sound signals. A sound processor, such as a DSP, can be used to process the sound signals to, for example, compensate for a user's hearing loss. The processor can be adapted to set its parameters based on an estimate of the direct reverberation ratio. The sound output device suitable for outputting the processed sound signal to the user can be a speaker or a cochlear implant.
[0053] It must be understood that the features of the methods described above and below can be the features of the computer programs, computer-readable media, and hearing devices described above and below, and vice versa.
[0054] These and other aspects of the invention will become apparent and will be elucidated with reference to the embodiments described below. Attached Figure Description
[0055] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings.
[0056] Figure 1 A hearing device according to an embodiment of the present invention is illustrated schematically.
[0057] Figure 2 A functional diagram of a hearing device is shown, illustrating a method for estimating the direct reverberation ratio of a sound signal according to an embodiment of the present invention.
[0058] Figure 3 and Figure 4 It shows having in Figure 2 A diagram of the start signal generated in the method.
[0059] Figure 5 It shows having in Figure 2 A graph of the start signal of the integral generated in the method.
[0060] Figure 6 Explanation is shown Figure 2 The performance graph of the method.
[0061] The reference numerals used in the figures and their meanings are listed in summary form in the reference numeral list. In principle, the same parts in the figures have the same reference numerals. Detailed Implementation
[0062] Figure 1A hearing device 10 in the form of a behind-the-ear device is illustrated schematically. It should be noted that the hearing device 10 is a particular embodiment, and the methods described herein can also be performed by other types of hearing devices (e.g., in-the-ear devices or wearable hearing devices).
[0063] The hearing device 10 includes a postauricular component 12 and a component 14 to be placed in the user's ear canal. Components 12 and 14 are connected by a tube 16. Component 12 includes a microphone 18, a sound processor 20, and a sound output device 22 (e.g., a speaker). The microphone 20 can acquire ambient sounds and generate sound signals, the sound processor 20 can amplify the sound signals, and the sound output device 22 can generate sound that is guided through the tube 16 and the in-ear component 14 into the user's ear canal.
[0064] The hearing device 10 may include a processor 24 adapted to adjust parameters of the sound processor 20, such as frequency-dependent amplification, frequency shift, and frequency compression. These parameters may be determined by a computer program running in the processor 24. For example, using a knob 26 on the hearing device 12, a user can select modifiers (e.g., bass, treble, noise suppression, dynamic volume, etc.), which affects the functionality of the sound processor 20. All these functions may be implemented as a computer program stored in the memory 28 of the hearing device 10, which may be executed by the processor 24.
[0065] Figure 2 It shows things like Figure 1 A functional diagram of a hearing device, such as a hearing aid. The boxes in the functional diagram may show the steps of the method as described herein and / or may show modules of the hearing device 10, such as software modules running in processor 24.
[0066] First, an audio signal 30 is acquired via microphone 18. For example, the audio signal can be recorded by hearing device 10 at a sampling frequency of 22050 Hz. The audio signal 30 can be buffered in time frames of 128 samples with 75% overlap.
[0067] Figure 3 and Figure 4 An audio signal 30 in the form of a speech signal is shown, which has a high direct reverberation ratio (8.7 dB). Figure 3 ) and a low direct reverberation ratio (-4.5 dB, Figure 4 The two figures illustrate the sound signal 30 relative to seconds in the time domain.
[0068] The audio signal 30, and specifically the time frames, can then be transformed from the time domain to the frequency domain using a Discrete Fourier Transform (e.g., a Fast Fourier Transform). A Hanning window and / or zero-padding can be applied before calculating the Discrete Fourier Transform.
[0069] The sound signal 30 is processed by the sound processor 20 to produce an output sound signal 32, which can then be output, for example, by the speaker 22. The operation of the sound processor 20 can be adjusted by means of a sound processor setting 34, which can be determined by the program 36 of the hearing device 10. These programs can also receive and evaluate the sound signal 30. For example, the program 36 can perform noise cancellation, reverberation cancellation, frequency-dependent amplification, frequency compression, beamforming, sound classification, self-voice detection, foreground / background classification, etc., by adjusting the sound processor 20 accordingly.
[0070] In particular, some or all of the programs 36 may receive the direct reverberation ratio 38 already determined based on the sound signal 30, and the programs 36 may additionally use the direct reverberation ratio 38 to determine the appropriate sound processor settings 34.
[0071] The direct reverberation ratio of 38 was determined as follows.
[0072] In the start determination box 32, the start signal 42 is determined based on the sound signal 30.
[0073] Typically, the audio signal 30 can be divided into time frames, which can be done before the discrete Fourier transform, and at least one energy value can be calculated for each time frame based on the audio signal 30. At least one start signal 42 can be determined based on the energy value, wherein the start value of the start signal 42 for the time frame is set to a positive value when the energy value of the time frame is higher than the energy value of the previous time frame by more than a threshold; and wherein otherwise, the start value is set to zero.
[0074] For example, for an audio signal 30 transformed to the frequency domain, the discrete Fourier transform box can be grouped into multiple sub-bands based on the ERB (Equivalent Rectangular Bandwidth) scale. For example, there might be 20 such sub-bands. Then, for each time frame and frequency sub-band E... k , f (k indicates the number of time frames and f indicates the frequency band; no sub-index f is needed for wideband cases) Calculate the power in dB or equivalent energy.
[0075] Therefore, the start signal 42 can be calculated.
[0076] Figure 3 and Figure 4Two different types of start signals are shown—wideband start signal 42a and frequency band start signal 42b. The start signals 42a and 42b in the corresponding figures correspond to the corresponding audio signal 30 in the top of the figures.
[0077] The broadband start signal 42a is determined based on the total power and / or energy of the audio signal 30 in time frames. A start is detected in frame k if the difference between the broadband power and / or energy of time frame k and time frame k-1 exceeds a given threshold. The broadband start signal 42a can be a binary feature, taking a value of 1 or 0 in each time frame. The value of the broadband start signal 42a at the kth time frame... It can be determined based on the following:
[0078]
[0079] Here, E k By applying all Ek from all subbands f, f The summation is performed to calculate the power and / or energy value of the k-th time frame.
[0080] The frequency band start signal 42b is determined based on the power and / or energy of the audio signal 30 within a specific frequency band's time frame. A frequency band can be determined by aggregating several sub-bands. For example, the 20 sub-bands mentioned above can be grouped into 4 frequency bands. The following illustrates how frequency bands can be divided.
[0081]
[0082]
[0083] The frequency bins of the Discrete Fourier Transform can also be grouped into frequency bands, and the power and / or energy of the bands can be calculated directly from the frequency bins. However, in many hearing devices, the sub-band energies mentioned above have been determined for other reasons.
[0084] For the band start signal 42c, the value for the i-th frequency range in the k-th time frame. The calculation rules can be
[0085]
[0086] exist Figure 3 and Figure 4 In the diagram, the band start signal 42c generated as a binary signal (i.e., having only values 0 and 1) is not shown, but the band start signal 42b is shown, wherein the value of the band start signal 42b is set to the strength of the start when a start is detected.
[0087] For the value of the frequency band start intensity 42b The calculation rules are almost identical to those for the band start signal 42c. However, in this case, whenever a start is detected, the power and / or energy difference is used as the value for that time frame.
[0088]
[0089] The starting strength of the broadband can also be determined in this way.
[0090] It is important to note that the frequency band threshold may differ from the threshold for the wideband start signal 42a. The thresholds may also differ for different frequency band start signals 42b and 42c.
[0091] In terms of the number of onsets, higher frequency ranges are generally more affected than lower frequency ranges. Therefore, reverberation typically does not exclusively reduce the number of onsets, but rather changes the onset distribution over time. Figure 3 and Figure 4 This can be directly seen from the start signal 42b shown in the diagram.
[0092] Figure 3 and Figure 4 The effect of reverberation on the intensity at the start of a frequency band is also shown. It can be seen that the overall intensity decreases with reverberation, and the highest frequency range is most affected.
[0093] Then, the start signal 42 is input into the machine learning algorithm 44. Typically, the direct reverberation ratio 38 can be determined by inputting at least one start signal 42 into the machine learning algorithm 44, which has been trained to generate the direct reverberation ratio 38 based on at least one start signal 42.
[0094] The machine learning algorithm 44 can be composed of several sub-blocks. The integrator 46 can determine the start signal 48 of the integration. The gradient determiner 50 can determine the gradient 52 of each start signal 48 of the integration, and the gradient 52 can be input into the regression model 54, which outputs a direct reverberation ratio 38.
[0095] Integrator 46 calculates the start signal 48 for integration based on each start signal 42 (specifically, based on start signals 42a, 42b, 42c). This is done by accumulating the start values of the corresponding start signals 42 over time. The value of the start signal 48 for integration over time frame k can be the sum of the values of the start signals 42 for time frames 0 to k.
[0096] Figure 5An example curve of a start signal 52 with several integrals for the same type of start signal 42 (e.g., wideband start signal 42a or band start signals 42a, 42b) is shown. In this figure, the number of time frames is depicted on the right. The curves have been determined for different known direct reverberation ratios 38 (DRR). It can be seen that when the direct reverberation ratio 38 is high, the total gradient and / or gradient 52 of the integral start signal 52 is also high.
[0097] The start signal 48 of the integration is input to the gradient determiner 50, which determines the gradient 52 for each start signal 48 of the integration. In particular, there exists a gradient 52 associated with each start signal 42a, 42b, 42c.
[0098] One or more gradients 52 can be determined by calculating the average gradient of the corresponding integral's start signal 48. This can be accomplished by determining the gradient at the far endpoint of the curve, such as... Figure 5 These gradients are depicted in the diagram. They can be averaged.
[0099] The gradient of a curve can also be determined using a state-space model.52 Using a state-space model, the gradient can be computed with less computational requirement because it avoids numerous divisions and / or inversions of large matrices. A state-space model can perform local line fitting on the accumulated start. Since the line is fully described by its gradient and intercept, the fitted parameters can directly represent these quantities. The intercept can be discarded, and the gradient can be preserved. A state-space model can be represented by a 2×2 matrix. Inversions to obtain the gradient can be avoided by using a pseudo-inverse matrix.
[0100] One or more gradients 52 are then input into the linear regression model 54 and / or used as features of the linear regression model 54. As described above, it can be assumed that gradient 52 indicates the reverberation associated with the start signals 42, 42a, 42b, 42c in the corresponding frequency bands, and reacts differently to changes in reverberation for different frequency bands. Therefore, gradient 52 is a good feature for machine learning algorithms.
[0101] The linear regression model 54 has been trained using gradients 52 extracted from sound signals 30 with different known direct reverberation ratios 38. The output of the linear regression model 54 is an estimate of the direct reverberation ratio 38.
[0102] The linear regression model 54 can have a weighted sum and / or coefficients of each gradient 52 input to it. The output of the linear regression model 54 (i.e., the estimated direct reverberation ratio 38) is the sum of these weighted sums and / or coefficients multiplied by the corresponding gradients 52. These weighted sums and / or coefficients are parameters that are adjusted during training.
[0103] It should be noted that one or more start signals 42 and / or gradients 52 can be input into another type of machine learning algorithm (e.g., artificial neural networks).
[0104] Figure 6 Graphs indicating the performance of direct reverberation ratio estimators 40 and 44 based on the azimuth of the sound source are shown. Specifically, graphs indicating the performance of direct reverberation ratio estimators for 36 directions of sound arrival are shown. It should be noted that the estimators are unaware of the direction of sound arrival. Values marked with circles refer to direct reverberation ratio estimates obtained using the method described herein from the left front microphone in a rear-ear hearing device.
[0105] The values marked with triangles refer to the direct reverberation ratio calculated based on the room impulse response recorded through the left front microphone and further measurements within the room. The values marked with triangles are affected by the directional pattern of the hearing device relative to the left rear azimuth, and the lateral values are affected by head shading. However, the same direct reverberation ratio should be determined on the same side of the body. It can be seen that the estimation is not affected by the directional pattern on the left rear side.
[0106] Although the invention has been described and illustrated in detail in the accompanying drawings and the foregoing description, such description and illustration should be considered illustrative or exemplary rather than restrictive; the invention is not limited to the disclosed embodiments. Other variations of the disclosed embodiments can be understood and implemented by those skilled in the art and those practicing the claimed invention through study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude multiple. A single processor or controller or other unit can perform the functions of several items cited in the claims. The fact that certain measures are cited in different dependent claims does not indicate that combinations of these measures cannot be advantageously used. No reference numerals in the claims should be construed as limiting the scope.
[0107] List of reference numerals
[0108] 10 hearing devices
[0109] 12 behind-ear parts
[0110] 14 Inner Ear Components 14
[0111] 16 tubes
[0112] 18 microphones
[0113] 20 sound processors
[0114] 22 audio output devices
[0115] 24 processors
[0116] 26 knobs
[0117] 28 memory
[0118] 30 sound signals
[0119] 32 Output sound signal
[0120] 34 Sound Processor Settings
[0121] 36 Hearing Device Programs
[0122] 38 direct reverberation ratio
[0123] 40 Start signal confirmed
[0124] 42 Start Signal
[0125] 42a Broadband Start Signal
[0126] 42b band start signal
[0127] 42C band start signal
[0128] 44 Machine Learning Algorithms
[0129] 46 Integrator
[0130] 48 points start signal
[0131] 50 Gradient Determiner
[0132] 52 gradients
[0133] 54 Regression Model
Claims
1. A method for estimating the direct reverberation ratio (38) of an acoustic signal (30), wherein, The direct reverberation ratio (38) indicates the ratio between the direct sound received from the sound source and the reverberant sound received from reflections in the environment of the sound source, the method comprising: Determine the first energy value of the sound signal (30) for the first time frame; If the difference between the first energy value of the first time frame and the second energy value of the previous second time frame is greater than a threshold, a positive value is assigned to the start value of the first time frame; otherwise, a zero value is assigned. The direct reverberation ratio (38) is determined by providing a start signal (42) including the start value to a machine learning algorithm (44), which has been trained to determine the direct reverberation ratio (38) based on the start signal.
2. The method according to claim 1, in, The start signal (42) is integrated over time to determine the gradient (52) of the start signal (42) and the gradient (52) is provided to the machine learning algorithm (44).
3. The method according to claim 2, in, The gradient (52) of the starting signal (48) of the integration is determined by the state-space model.
4. The method according to any one of the preceding claims, in, The machine learning algorithm (44) includes a linear regression model (54).
5. The method according to claim 1, in, A broadband energy value is determined for the first time frame, the broadband energy value indicating the energy of the sound signal (30) in the first time frame; Wherein, when the broadband energy value of the first time frame is higher than the broadband energy value of the previous second time frame by more than a broadband threshold, the broadband start signal (42a) is determined by setting the broadband start value of the broadband start signal (42a) for the first time frame to a positive value.
6. The method according to claim 1, in, A frequency band energy value is determined for the first time frame, the frequency band energy value indicating the energy of the sound signal (30) in the frequency band of the first time frame; Wherein, when the frequency band energy value of the first time frame is higher than the frequency band energy value of the previous second time frame by more than a frequency band threshold, the frequency band start signal (42b) is determined by setting the frequency band start value of the frequency band start signal (42b) for the first time frame to a positive value.
7. The method according to claim 6, in, The sound signal (30) is divided into multiple frequency bands, and a frequency band start signal (42b) is determined for each frequency band.
8. The method according to claim 6 or 7, in, The frequency band threshold is different from the broadband threshold; and / or The frequency band thresholds are different for different frequency bands.
9. The method according to claim 1, in, The initial value is set to the positive value of 1; or The positive value is the difference between the energy value in the first time frame and the energy value in the previous second time frame.
10. The method according to claim 9, in, The broadband start signal (42a) is determined when the positive value is set to 1; Among them, when the positive value is set to 1, multiple first frequency band start signals (42c) for multiple frequency bands are determined; Among them, a plurality of second frequency band start signals (42b) for the plurality of frequency bands are determined when the positive value is set to the difference between the energy value in the first time frame and the energy value in the previous second time frame; The broadband start signal (42a), the first frequency band start signal (42c), and the second frequency band start signal (42b) are input into the machine learning algorithm.
11. A method for operating a hearing device (10), the method comprising: A sound signal (30) is generated using the microphone (18) of the hearing device (10); According to any one of the preceding claims, the direct reverberation ratio (38) of the sound signal (30) is estimated; The sound signal (30) is processed using the direct reverberation ratio (38) to compensate for the hearing loss of the user of the hearing device (10); The processed audio signal (32) is output to the user.
12. The method according to claim 11, in, The direct reverberation ratio (38) is used for at least one of the following: Noise cancellation Reverb elimination Frequency-dependent amplification Frequency compression, Beamforming Sound classification Self-voice detection Foreground / Background Classification.
13. A computer program for estimating the direct reverberation ratio of an acoustic signal (30), said computer program being adapted, when executed by a processor, to perform the steps of the method described in any one of the preceding claims.
14. A computer-readable medium (28) storing the computer program according to claim 13.
15. A hearing device (10) adapted to perform the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Determination of Room Reverberation for Signal Enhancement
US20170303053A1
Apparatus and method for determining a measure for a perceived level of reverberation, audio processor and method for processing a signal
CN103430574A
Determination of room reverberation for signal enhancement
CN106688247A