Methods and apparatus for processing audio signals in a frequency-selective manner with low latency

By using multi-level decomposition filter banks and frequency domain prediction technology, the conflict between low latency and high frequency resolution in hearing devices is resolved, achieving high frequency resolution and high-quality audio signal processing.

CN115379366BActive Publication Date: 2026-03-06SIVANTOS PTE LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-17
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Existing hearing devices struggle to balance low latency and high frequency resolution when processing audio signals, resulting in sound quality distortion and poor processing performance.

Method used

A multi-stage decomposition filter bank combined with frequency domain prediction technology is used to split the input audio signal into frequencies, and the delay difference is compensated by a predictor, especially for fine frequency division and processing of the low-frequency part in the frequency domain.

Benefits of technology

It achieves high frequency resolution and good sound quality across a portion of the sound spectrum, reduces the negative impact of latency on sound quality, and improves the effectiveness of signal processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115379366B_ABST
    Figure CN115379366B_ABST
Patent Text Reader

Abstract

A method for processing an input audio signal is provided. Here, in a first frequency split, the input audio signal is divided into multiple first frequency bands using a first decomposition filter bank. In at least one additional frequency split, the first frequency band of a first subgroup of the first frequency band is divided into multiple sub-frequency bands using at least one additional decomposition filter bank. The input audio signal divided into first frequency bands or sub-frequency bands is processed in a frequency-selective manner, particularly amplified. The input audio signal divided into first frequency bands or sub-frequency bands and processed in a frequency-selective manner is combined into an output audio signal. According to the method, prediction is applied to the first frequency band of the first subgroup and / or the sub-frequency bands derived therefrom to compensate for delay differences between the first frequency band and sub-frequency bands caused by the additional frequency split or each additional frequency split. The device constructed for performing the method is particularly formed from a hearing device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method and apparatus for processing (input) audio signals. Such a method and apparatus are known from EP 2 124 335 B1. The apparatus is, in particular, a hearing device. Background Technology

[0002] Electronic devices that support the hearing of individuals wearing hearing aids (hereinafter referred to as "wearers" or "users") are generally called "hearing devices" or "hearing apparatuses." In particular, the present invention relates to hearing devices configured to fully or partially compensate for the hearing loss of users with hearing impairments. Such hearing devices are also referred to as "hearing aids." Furthermore, there are hearing devices that protect or improve the hearing of users with normal hearing, for example, to improve speech comprehension in complex hearing situations. Additionally, headphones or other sound reproduction devices also fall under the category of hearing devices, which insert ambient noise into another audio signal (e.g., music or telephone calls) or reduce the perception of ambient noise through active noise suppression.

[0003] Generally speaking, hearing devices, specifically hearing aids, are typically constructed to be worn on a user's head, particularly in or on their ear, especially as behind-the-ear (BTE) devices or in-the-ear (ITE) devices. Internally, hearing devices typically have at least one (sound-to-electrical) input converter, a signal processor, and an output converter. When the hearing device is in operation, the input converter receives airborne sound from the environment and converts it into an input audio signal (i.e., an electrical signal conveying information about the ambient sound). The signal processor processes the input audio signal (i.e., modifies it according to its sound information) to support the user's hearing, particularly to compensate for hearing loss. The signal processing unit outputs the corresponding processed audio signal (also called the "output audio signal" or "modified sound signal") to the output converter.

[0004] In most cases, the output converter is constructed as an electro-acoustic converter, which converts the (electrical) output audio signal back into airborne sound, wherein this (modified relative to ambient sound) airborne sound is output into the user's ear canal. In the case of hearing devices worn behind the ear, the output converter, also known as a "receiver," is typically integrated into the housing of the hearing device outside the ear. In this case, the sound output by the output converter is guided into the user's ear canal via a sound tube. Alternatively, the output converter can also be placed inside the ear canal, and thus can be placed outside the housing worn behind the ear. Such hearing devices (receiver incanal) are also called RIC devices. Hearing devices worn in the ear (completely in canal) are also referred to as such, as these hearing devices are designed to be so small that they do not protrude out of the ear canal.

[0005] In other structural forms, the output transducer can also be configured as an electromechanical transducer that converts the output audio signal into solid-borne sound (Körperschall) (vibration), which, for example, is output to the user's skull. Furthermore, there are implantable hearing devices, particularly cochlear implants, and hearing devices in which the output transducer directly stimulates the user's auditory nerve.

[0006] For signal processing, the input audio signal in a hearing device is typically divided into multiple frequency bands using a decomposition filter bank. In other words, the input audio signal is converted into multiple sub-band signals, which are then guided separately in the frequency channels and processed, in particular, amplified, in a specific manner. The processed sub-band signals are then recombined into an output audio signal (including all frequency components) using a synthesis filter bank.

[0007] One crucial aspect of audio signal processing, particularly in hearing aids, is latency—the time delay in the audio signal caused by processing. Ideally, latency should be less than 10 milliseconds (ms), as greater delays significantly impact the listening experience for human users. On the other hand, high frequency resolution is beneficial for common signal processing functions such as noise reduction or dynamic compression; that is, dividing the audio signal into numerous frequency bands, each with a small bandwidth, is meaningful.

[0008] However, these two requirements conflict with each other because the product of time resolution and frequency resolution is constant (Kupfmüller's uncertainty principle). Therefore, in practice, the technically meaningful minimum bandwidth is limited by the maximum tolerable delay, which in some cases makes satisfactory signal processing difficult or even impossible.

[0009] Therefore, for example, to achieve good noise reduction for spoken (or voiced) speech, a frequency interval between adjacent frequency bands corresponding to at least half of the fundamental frequency is desirable. However, even for female voices, whose fundamental frequency is typically between 200 Hz and 300 Hz, a desired frequency interval of 100 Hz to 150 Hz cannot be achieved because this interval would be associated with excessive delay. Therefore, in common hearing devices, a frequency resolution of 200 Hz to 500 Hz is generally achieved as an acceptable, but not entirely satisfactory, trade-off between the highest possible frequency resolution and the lowest possible delay.

[0010] To address this problem, filter banks are sometimes used where the frequency bands have non-uniform frequency intervals, meaning the intervals increase continuously or abruptly with increasing frequency. Therefore, EP 2 124 335 B1 discloses a two-stage decomposition filter bank device for hearing devices, in which the audio signal to be processed is divided into four first frequency bands by a first filter bank, and then further divided into 24 second frequency bands by a second filter bank. The lower twelve of these 24 second frequency bands have significantly smaller frequency intervals and smaller bandwidths compared to the upper twelve frequency bands.

[0011] The non-uniform frequency resolution of known decomposed filter bank devices reduces the drawbacks caused by high frequency splitting because increased delays only occur in a portion of the spectrum. However, this advantage comes at the cost of degraded sound quality, as the output audio signal is distorted due to the large delays and different group run times in the low-frequency and high-frequency channels. Furthermore, in different forms of filter banks, each subband generally has a different bandwidth. However, potential undersampling must be directed towards the band with the highest bandwidth. This results in relatively inefficient signal processing.

[0012] On the other hand, EP 3 197 181 A1 describes a method and apparatus for processing audio signals in a hearing device, wherein multiple signal blocks in the time domain are formed from the input audio signal. To reduce delay, at least some of these time blocks are predicted, i.e., extrapolated to future signal changes in these time blocks. The predicted time blocks are then divided into frequency bands by a filter bank, thus transforming them into the frequency domain. However, this known method also suffers from a significant impact on sound quality due to the prediction involved. Summary of the Invention

[0013] The technical problem to be solved by this invention is to enable frequency-selective processing of audio signals with low latency and high quality (sound quality). In particular, it is to enable high frequency resolution across a portion of the audible sound spectrum.

[0014] According to the present invention, the aforementioned technical problems are solved by the features of the present invention in terms of methods for processing audio signals (frequency-specific). According to the present invention, the aforementioned technical problems are solved by the features of the present invention in terms of apparatuses for processing audio signals (frequency-specific). The explanatory portions in the following description are themselves considered to be advantageous designs and extensions of the present invention that are inventive.

[0015] In the process of the method, in order to process the input audio signal, particularly in hearing devices, the input audio signal is first divided into multiple first frequency bands (first frequency splits) in the spectrum using a first decomposition filter bank. In at least one further frequency split, a first subgroup ( subset) of the first frequency band is divided into sub-frequency bands (narrower than the first frequency band) using at least one additional decomposition filter bank. The input audio signal, divided into the first frequency bands and, if necessary, sub-frequency bands, is processed, particularly amplified, in a frequency-selective manner. Then, the input audio signal, divided into the first frequency bands and, if necessary, sub-frequency bands, and processed in a frequency-selective manner, is recombined into an output audio signal.

[0016] According to the invention, a prediction is now applied to the finer frequency division portion of the input audio signal (preferably only to this portion of the input audio signal), which compensates for the delay caused by further frequency division or each subsequent frequency division, i.e., completely eliminating or at least reducing such delay. In other words, the delay difference between frequency bands and sub-bands caused by further frequency division or each subsequent frequency division is compensated for by making a prediction. In particular, the delay of the finer frequency division portion of the input audio signal is adapted to the smaller delay of the coarser frequency division portion of the input audio signal.

[0017] Unlike the method known from EP 3 197 181 A1, prediction is applied in the frequency domain here. Within the scope of this invention, there are several variations regarding the time point or location for applying prediction in the frequency domain. Therefore, prediction is applied directly to sub-bands and / or to first frequency bands from which sub-bands are derived. Here, prediction can be performed accordingly before or after signal processing, or between two of several possible processing steps. Finally, within the scope of this invention, prediction can also be performed in multiple consecutive prediction steps.

[0018] The method described above enables particularly fine frequency splitting over a portion of the sound spectrum, where prediction is performed while avoiding or at least reducing distortion in the output signal that typically accompanies uneven frequency splitting. However, compared to the method known from EP 3 197 181 A1, the adverse effects of prediction on sound quality are also reduced because prediction is applied only to a portion of the sound spectrum. Therefore, high frequency resolution and exceptionally good sound quality are achieved overall over a portion of the sound spectrum.

[0019] In a preferred embodiment of the method, frequency splitting is performed in a two-stage manner. Here, using a second decomposition filter bank, the first subgroup (set) of the first frequency band is further subdivided into second frequency bands (i.e., second-stage subbands) in the second frequency splitting. That is, each first frequency band of the first subgroup is further divided into multiple such second frequency bands. Here, prediction is applied to either the first or second frequency band of the first subgroup to compensate for the delay caused by the second frequency splitting.

[0020] Optionally, the basic idea of ​​the method according to the invention, namely, to perform finer frequency splitting on the spectral portion of the input audio signal and to predict the more finely divided spectral range to compensate for the delay caused by the finer frequency splitting, is extended to n-level frequency splitting (where n = 3, 4, 5, ...). Here, generally speaking, the subgroup of the i-th frequency band (where i = 2, 3, 4, ...) is divided into a further narrower (i+1)-th frequency band. Accordingly, the multiple frequency-divided portions of the input audio signal are predicted in the frequency domain, such that the delay caused by the multiple frequency splits is compensated accordingly.

[0021] Therefore, in the three-level implementation of this method principle, by means of a third decomposition filter bank, each of the second frequency bands of the second frequency band subgroup is divided into multiple third frequency bands (i.e., third-level sub-bands) in the third frequency split. Here, prediction is applied to the third frequency bands, and / or to the second frequency bands from which the third frequency bands are derived, and / or to the first frequency bands from which the third frequency bands are derived, so as to compensate for the delay caused by the second and third frequency splits.

[0022] It is preferable to select a first subgroup of the first frequency band such that it covers a continuous low-frequency range of the sound spectrum, particularly the lower 2 to 3 kHz. In other words, the first subgroup of the first frequency band is preferably formed by a plurality of first frequency bands whose center frequencies are directly adjacent and which include the lowest first frequency band. This is particularly advantageous for processing audio signals containing human speech. This is because, on the one hand, the sound component of speech noise, especially for spoken speech, dominates in this low-frequency range, and on the other hand, the frequency resolution of human hearing is particularly high at low frequencies.

[0023] In principle, the method can be used in conventional multi-stage filter banks, as known for example from EP 2 124 335 B1. However, it is preferable to use additional decomposed filter banks, or each additional decomposed filter bank, that operates only on the first subgroup of the first frequency band. Conversely, it is preferable to process, and in particular amplify, the second subgroup of the first frequency band in a frequency-selective manner without further frequency splitting. This results in a particularly low overall delay.

[0024] In an advantageous embodiment of the invention, particularly efficient frequency splitting and processing of the input audio signal is achieved by having a uniform first bandwidth, i.e., the same first bandwidth for all first frequency bands. For the same reason (besides alternatives), it is also preferable to design the i-th frequency bands (where i = 2, 3, ...) such that these i-th frequency bands correspondingly have a uniform i-th bandwidth, i.e., the same i-th bandwidth for all i-th frequency bands. Here, the first bandwidth is specifically an integer multiple of the second bandwidth; the second bandwidth may be an integer multiple of the third bandwidth, and so on.

[0025] In principle, within the scope of this invention, linear prediction can be performed on portions of the input audio signal that are more finely divided into frequencies. However, it is preferable that the prediction applied to the first frequency band of the first subgroup or the subband derived therefrom is a nonlinear prediction.

[0026] In a particularly advantageous embodiment of the invention, one or more adaptive prediction algorithms are used during the operation of the method, i.e., during signal processing. Unlike non-adaptive (static) pre-configured or trained prediction algorithms during the operation of the method, adaptive prediction algorithms are highly flexible and resource-efficient, and are therefore particularly suitable for use in hearing devices.

[0027] In suitable embodiments of the invention, for the execution of predictions, in particular at least one Hammerstein model, a recurrent neural network (Netzwerk), and / or an echo-state network (Netzwerk) is used.

[0028] To further reduce the overall negative impact of uneven frequency splitting and prediction of the input audio signal on the sound quality of the output signal, in an extension of the invention, the methods described above are used only occasionally when they offer particular advantages, especially when processing sounds containing spoken speech. For this purpose, the input audio signal is analyzed (in a cross-band or band-specific manner) based on the presence or absence of spoken speech. Here, only when spoken speech is identified in the input audio signal is further frequency splitting, or each further frequency splitting, performed in at least one of the first frequency bands or sub-bands, on the signal path leading to the output audio signal, thereby also performing prediction. However, alternatively, further frequency splitting, or each further frequency splitting and / or prediction can continue in the context of signal processing, even when no spoken speech is present, without affecting the output audio signal.

[0029] As an addition or alternative, for the same purpose, the accuracy (reliability) of the prediction is determined (either cross-band or band-specific). Here, only when the accuracy of the prediction there meets a pre-given standard, particularly exceeding a pre-given threshold, is a more refined frequency split (i.e., a derived sub-band) performed on a portion of the input audio signal for the signal path leading to the output audio signal, thereby also performing prediction. However, alternatively, without affecting the output audio signal, in the context of signal processing, and when the accuracy of the prediction is insufficient, further frequency splitting or prediction may continue with each further frequency split and / or prediction.

[0030] The device according to the invention is particularly a hearing device having one of the structural forms described at the beginning, preferably a hearing aid.

[0031] The device includes a decomposition filter bank arrangement having a first decomposition filter bank and at least one additional decomposition filter bank. The first decomposition filter bank is configured to divide the input audio signal into a plurality of first frequency bands. At least one additional decomposition filter bank is connected after the first decomposition filter bank and is configured to divide each first frequency band of a first subgroup of the first frequency bands into a plurality of sub-frequency bands. As described above, in addition to the second decomposition filter bank, the decomposition filter bank arrangement may optionally include a third decomposition filter bank connected downstream, which further and more finely divides the subgroups of the second frequency bands into third frequency bands, and may further include one or more decomposition filter banks further connected downstream if necessary.

[0032] In addition, the device includes: a signal processing unit for selectively processing, particularly amplifying, an input audio signal divided into a first frequency band or a sub-frequency band; and a synthesis filter bank connected downstream of the signal processing unit, configured to combine the input audio signal, which has been divided into a first frequency band, or, if necessary, a sub-frequency band, and processed selectively, into an output audio signal.

[0033] According to the invention, the device includes at least one predictor configured to apply predictions to a first frequency band and / or sub-frequency bands derived therefrom of a first subgroup to compensate for delay differences between the first frequency band and sub-frequency bands resulting from further frequency splits or each subsequent frequency split.

[0034] Preferably, the signal processing unit is implemented in the digital signal processor of the hearing device. Within the scope of this invention, the signal processing unit can be implemented in the form of (non-programmable) electronic circuitry. Here, the signal processor is constructed, for example, as an ASIC or includes an ASIC. Alternatively, the signal processing unit is implemented in software. In this case, the signal processor is formed by programmable electronic components. Again, as an alternative, the signal processing unit is formed by a combination of non-programmable circuitry and software. Here, the signal processor is formed by a hybrid chip including at least one programmable component and at least one non-programmable component.

[0035] The synthesized filter bank device is preferably configured to be mirror-symmetric to the decomposed filter bank device, thus including a corresponding pair for each decomposed filter bank. Specifically, the synthesized filter bank device includes: a second synthesized filter bank that combines the second frequency band into the first frequency band after signal processing; and a first synthesized filter bank that combines the first frequency band into an output signal. In embodiments where the decomposed filter bank device includes more than two decomposed filter banks, the synthesized filter bank preferably also includes a corresponding plurality of synthesized filter banks.

[0036] Generally speaking, the device according to the invention is set up and configured to automatically perform the method according to the invention described above. Here, the design schemes and extensions of the method described above correspond to the corresponding design schemes and extensions of the device. Therefore, the description of the necessary and optional features of the method and its corresponding effects and advantages can be applied to the device, and vice versa.

[0037] Therefore, in the preferred design, the device is configured as follows:

[0038] • The first subgroup of the first frequency band is formed by a plurality of first frequency bands within the first frequency band, wherein the center frequencies of these first frequency bands are directly adjacent and include the lowest first frequency band.

[0039] • The second subgroup of the first frequency band is directly fed to the signal processing unit so that the second subgroup of the first frequency band can be processed in a frequency-selective manner without further frequency splitting.

[0040] • The first frequency band and / or the i-th level (where i = 2, 3, 4, ...) sub-bands have a uniform first or i-th bandwidth, wherein the first bandwidth is in particular an integer multiple of the second bandwidth, and so on.

[0041] • At least one predictor is nonlinear, and / or

[0042] • At least one predictor is adaptive during runtime, i.e., during signal processing.

[0043] Optionally, the device includes: a speech recognition module configured to analyze the input audio signal for the presence or absence of spoken speech; and a switching device, also known as a "signal switch," configured to activate an additional decomposition filter bank or each additional decomposition filter bank only when the speech recognition module identifies the presence of spoken speech in the input audio signal.

[0044] Alternatively or additionally, the device includes a switching device (signal switcher) that is the same as or different from the switching device described above, configured to activate and deactivate additional decomposition filter banks or each additional decomposition filter bank based on the accuracy (reliability) of the prediction. Within the scope of the invention, in principle, the accuracy of the prediction can be determined by the switching device itself by analyzing the sub-band signals correspondingly guided in the first frequency band of the first subgroup. However, it is preferable that the predictor or each predictor determines a characteristic parameter characterizing the accuracy of the prediction and outputs it to the switching device, which activates and deactivates the second decomposition filter bank based on this characteristic parameter. In particular, the so-called "prediction gain" is used as a characteristic parameter for the accuracy of the prediction.

[0045] The decomposition filter bank device, the synthesis filter bank device, the predictor or each predictor, and (if present) the speech recognition module and / or the switching device or each switching device are preferably integrated (in the form of (non-programmable hardware) and / or software) into the signal processor of the device. In particular, within the scope of the invention, the switching device or each switching device may also be a software module. Attached Figure Description

[0046] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings. Wherein:

[0047] Figure 1 The diagram illustrates a hearing aid-type device that can be worn behind a user's ear.

[0048] Figure 2 A schematic block diagram is shown. Figure 1 The structure of signal processing in hearing devices, and

[0049] Figure 3 and 4 According to Figure 2 The illustration shows two alternative implementations of the hearing device.

[0050] In all the accompanying drawings, corresponding parts and parameters are always given the same reference numerals. Detailed Implementation

[0051] As an example of a device for processing audio signals according to the present invention, Figure 1 Hearing aid 2 is shown, which is a hearing device configured to support the hearing of a user with hearing loss. In the example shown here, hearing aid 2 is a BTE hearing aid that can be worn behind the user's ear.

[0052] The hearing aid 2 includes, within the housing 4, at least one microphone 6 as an input converter and an earpiece 8 as an output converter. Furthermore, the hearing aid 2 includes a battery 10 and, particularly, a digital, signal processor 12. Preferably, the signal processor 12 includes not only programmable subunits (e.g., microprocessors) but also non-programmable subunits (e.g., ASICs).

[0053] The battery 10 supplies the power supply voltage U to the signal processor 12.

[0054] During normal operation of the hearing aid 2, the microphone 6 receives airborne sound from the environment surrounding the hearing aid 2. The microphone 6 converts this sound into an (input) audio signal I, which contains information about the received sound. Inside the hearing aid 2, the input audio signal I is fed to the signal processor 12, which modifies the input audio signal I to support the user's hearing.

[0055] The signal processor 12 outputs an audio signal O to the earpiece 8. The audio signal O contains information about the processed and thus modified sound.

[0056] The earpiece 8 converts the output audio signal O into a modified airborne sound. This modified airborne sound is transmitted to the user's ear canal via the sound channel 14 connecting the earpiece 8 and the housing 4, and via a flexible sound tube (not explicitly shown) connecting the tip 16 and the earplug inserted into the user's ear canal.

[0057] exist Figure 2 The functional structure of the signal processor 12 is shown in more detail below.

[0058] First, the input audio signal I received by the microphone 6 is digitized via an analog-to-digital converter integrated in or connected upstream of the signal processor 12, in a manner not shown in detail. Inside the signal processor 12, the digitized input audio signal I is first fed to the decomposition filter bank device 20. Figure 2 In the example shown, the decomposition filter bank device 20 includes a first decomposition filter bank 22 and a second decomposition filter bank 24 connected downstream of the first decomposition filter bank 22.

[0059] With the aid of the first decomposition filter bank 22, the input audio signal I is divided into multiple first frequency bands 26, i.e., first frequency channels, in the first frequency split, which respectively guide the sub-band signals of the input audio signal I. Figure 2Only four first frequency bands 26 are simply shown. In a meaningful practical implementation of the invention, the first decomposition filter bank 22 divides the input audio signal I into, for example, 32 first frequency bands 26. The frequency bands 26 have a uniform (first) bandwidth of, for example, 500 Hz and a uniform spectral spacing of 250 Hz.

[0060] The second decomposition filter bank 24 operates only on the (first) subgroup 28 of frequency band 26, which covers a range of 2 to 3 kHz at the low-frequency edge of the sound spectrum. Here, subgroup 28 comprises multiple adjacent frequency bands 26, including the lowest (i.e., lowest frequency) first frequency band 26. Figure 2 In the example shown, subgroup 28 exemplarily includes the two lower of a total of four frequency bands 26; in a real implementation, subgroup 28 may include, for example, the 12 lower of a total of 32 first frequency bands 26.

[0061] In the second frequency split, the second decomposition filter bank 24 splits each of the first frequency bands 26 of the sub-group 28 into multiple (according to) Figure 2 (Example two) Second frequency bands 30. Frequency band 30 has, for example, a uniform (second) bandwidth of 125 Hz and a uniform spectral spacing of 62.5 Hz.

[0062] The second subgroup 32 of the frequency band 26, which includes high frequencies that do not belong to subgroup 28, is passed through the second decomposition filter group 24, so that it is not subjected to a second (and more refined) frequency split.

[0063] In signal processing unit 34, the high-frequency band 26 and the sub-band signals of each of the bands 30 of subgroup 32 are processed (i.e., modified using signal technology). During this processing, in particular, the sub-band signals of each band 26 and each band 30 of subgroup 32 are amplified according to individual (i.e., pre-given in a frequency-specific manner) amplification factors. In the sense of effective signal processing, according to... Figure 2 In the example, the signal processing unit 34 includes a high-frequency band 26 for subgroup 32 and two subunits 36 and 38 for band 30, wherein subunits 36 and 38 are specifically designed for different bandwidths of the fed bands 26 and 30, respectively.

[0064] The synthesis filter bank device 40 combines the high-frequency band 26 and the sub-band signal of band 30 of the processed subgroup 32 into an output audio signal O. The synthesis filter bank device 40 is designed to be mirror-symmetrical to the decomposition filter bank device 20. It therefore includes a second synthesis filter bank 42 and a first synthesis filter bank 44. The second synthesis filter bank 42 combines the second band 30 into the first band 26 of subgroup 28, and the first synthesis filter bank 44 combines the first subgroup 28 and the first frequency channel 26 of subgroup 32 into the output audio signal O.

[0065] Due to the finer frequency splitting achieved by the decomposition filter bank 24, a delay difference is caused in the low-frequency subband signal of the frequency band 26 of the subgroup 28 compared to the high-frequency subband signal of the frequency band 26 of the subgroup 32. This difference will result in distortion of the output audio signal O without further measures.

[0066] To compensate for this delay difference (i.e., to completely eliminate or at least reduce it), a predictor 46 is connected in the signal path of frequency band 30. Predictor 46 is preferably configured to be nonlinear and to be a continuously adaptive predictor during operation of the hearing aid 2, particularly configured as a Hammerstein model. Predictor 46 has specifically adapted parameters for each fed frequency band 30.

[0067] Unlike the method known from EP 3 197 181 A1, prediction in hearing aid 2 is performed in the frequency domain, and prediction is applied only to the low-frequency portion of the sound spectrum, which has been more finely divided into frequency groups. However, within the frequency domain, i.e., between the first decomposition filter group 22 and the first synthesis filter group 44, the predictor 46 can be positioned differently. According to... Figure 2 In the example, predictor 46 is connected between the second decomposition filter bank 24 and subunit 38 of signal processing unit 34. Furthermore, in Figure 2 The figure shows three alternative positions for the predictor (marked here by reference numeral 46'), namely,

[0068] • Between the first decomposition filter group 22 and the second decomposition filter group 24

[0069] • Between subunit 38 of signal processing unit 34 and the second synthesis filter bank 42, and

[0070] • Between the second synthesis filter group 42 and the first synthesis filter group 44.

[0071] exist Figure 2 In one possible variation of the embodiment shown, the hearing aid 2 includes a plurality of interconnected predictors 46, 46', which are particularly arranged in Figure 2The multiple locations given in the text are used to compensate for a portion of the delay difference described above.

[0072] A digital-to-analog converter (not shown in detail) integrated in or connected downstream of the signal processor 12 converts the output audio signal O from the first synthesis filter bank 44 back into an analog signal and feeds it to the earpiece 8 for output to the user of the hearing aid 2.

[0073] exist Figure 3 An alternative embodiment of the hearing aid 2 is shown, wherein the decomposition filter bank device 20 and the synthesis filter bank device 40 are each constructed as three stages. In this embodiment, in addition to the first decomposition filter bank 22 and the second decomposition filter bank 24, the decomposition filter bank device 20 also includes a third decomposition filter bank 50 acting on a (first) subgroup 52 of the frequency band 30. The subgroup 52 further covers the low-frequency portion of the sound spectrum that is generally spanned by the frequency band 30. For example, the subgroup 52 is particularly (in accordance with...) Figure 3 In the example, it includes the two lower bands of a total of four frequency bands 30; in the actual implementation, subgroup 52 includes, for example, the six lower bands of a total of 12 second frequency bands 30.

[0074] The third decomposition filter bank 50 further and more finely splits each of the second frequency bands 30 of sub-group 52 into multiple (according to...) Figure 3 Exemplary example: two) third frequency bands 54 (third frequency splits). Frequency band 54 has, for example, a uniform (third) bandwidth of 62.5 Hz and a uniform spectral spacing of 31.25 Hz.

[0075] The second subgroup 56 of the frequency band 30, which includes high frequencies that do not belong to subgroup 52, is passed through the third decomposition filter group 50, so that it is not subjected to a third frequency split.

[0076] According to Figure 3 In the embodiment of the hearing aid 2, the subunit 38 of the signal processing unit 34 processes only the sub-band signals of the high-frequency band 30 of the second subgroup 56. To process the sub-band signals of the band 54, particularly to amplify them in a frequency-selective manner, according to... Figure 3 The signal processing unit 34 additionally includes another subunit 58, which is designed for the bandwidth of the frequency band 54.

[0077] According to Figure 3 In this embodiment, the synthesis filter bank device 40 is also configured to be mirror-symmetrical to the decomposition filter bank device 20. Therefore, in addition to the first synthesis filter bank 44 and the second synthesis filter bank 42, it also includes a third synthesis filter bank 60, which, after signal processing, combines the third frequency band 54 into the second frequency band 30 of the subgroup 52.

[0078] According to Figure 3 In the implementation of the hearing aid 2, the predictor 46 also operates only on the high-frequency sub-band signals of the second frequency band 30 of the second subgroup 56. To predict the low-frequency sub-band signals of the second frequency band 30 and the third frequency band 54 of the first subgroup 52, according to... Figure 3 The hearing aid 2 includes another predictor 62. Predictor 62 preferably has the same type as predictor 46, but is designed to compensate for the delay difference caused by the splitting of the second and third frequencies for the subband signal of the first frequency band 26 of the subgroup 32.

[0079] The predictor 62 can also be arranged at different locations between the second decomposition filter bank 24 and the second synthesis filter bank 42. Furthermore, according to... Figure 3 In a variation of the implementation, multiple predictors 62 may be connected in series, each compensating for a portion of the delay difference. In another variation, predictors 46 and 62 are connected in series. Here, predictor 46 is arranged between the first decomposition filter group 22 and the second decomposition filter group 24, or between the second synthesis filter group 42 and the first synthesis filter group 44. In this case, predictor 62 is designed to compensate only for the delay difference caused by the third frequency split.

[0080] Figure 4 Another embodiment of hearing aid 2 is shown, which is basically corresponding to the one described in the figure. Figure 2 The implementation method. However, compared with... Figure 2 The implementation method is different. Here, the first frequency band 26 of the low frequency of subgroup 28 is split into two frequencies only when the input audio signal I contains spoken speech (i.e., the sound of speaking or singing).

[0081] To this end, a speech recognition module 64 is implemented in the signal processor 12. The speech recognition module 64 identifies the presence of spoken speech by analyzing the input audio signal I, which is decomposed into a first frequency band 26, particularly the low-frequency portion therein. In the example shown, the frequency band 26 of subgroup 28 is fed to the speech recognition module 64 as an input parameter. Here, the speech recognition module 64 identifies the presence of spoken speech, particularly by the presence of a distinct fundamental frequency and / or the appearance of dominant frequencies (formants) characterizing spoken sound. When spoken speech is identified in the input audio signal I, the speech recognition module 64 outputs a control signal S1.

[0082] To perform the second frequency split only when spoken speech is recognized, a signal switcher 66 is connected in the signal path of frequency band 26 of subgroup 28. The signal switcher 66 guides the sub-band signal of frequency band 26 of subgroup 28 to the second decomposition filter bank 24 or the sub-unit 36 ​​of the data processing unit 34 according to the control signal S1. When the control signal S1 is applied (and therefore when spoken speech is recognized in the input audio signal I), the signal switcher 66 guides the sub-band signal of frequency band 26 of subgroup 28 to the second decomposition filter bank 24. In this case, Figure 4 The function of hearing aid 2 in the middle corresponds to that in Figure 2 The implementation shown is different. Conversely, when control signal S1 is not applied to signal switcher 66 (so speech recognition module 64 does not recognize spoken speech in input audio signal I), signal switcher 66 directly guides the sub-band signals of frequency band 26 of subgroup 28 to sub-unit 36 ​​of data processing unit 34. In this case, all sub-band signals of the first frequency band 26 are processed, particularly amplified in a frequency-specific manner, without further frequency splitting. Thus, prediction is also not performed.

[0083] exist Figure 4 In an alternative implementation of the hearing aid 2, the second frequency split is activated not based on the recognition of spoken speech, but on the accuracy (reliability) of the prediction. Here, the predictor 46 outputs a characteristic parameter Q characterizing the accuracy of the prediction, specifically the so-called "predictor gain," which is expressed in decibels as a percentage of the variance relative to the prediction error. The variance of the input signal of the predictor 46 To provide:

[0084]

[0085] When (as in) Figure 4 When multiple subband signals (as shown) are fed to the predictor as input signals, the characteristic parameter Q is calculated, for example, from the average, minimum, or maximum value of the predictor gain for each frequency band. Alternatively, the predictor gain of the subband signal selected as a reference is used as the characteristic parameter Q. In all these cases, the larger the value of the characteristic parameter Q, the more accurately the predictor 46 can predict changes in the fed subband signals.

[0086] Implemented in signal processor 12 (and in Figure 4 The evaluation module 68 (shown as dashed in the image) compares the feature parameter Q with a pre-defined threshold. If the feature parameter Q exceeds the threshold, the evaluation module 68 outputs a control signal S2 instead of the control signal S1, and feeds the control signal S2 to the signal switcher 66.

[0087] When control signal S2 is applied (and therefore, if the prediction accuracy is sufficient), signal switcher 66 directs the sub-band signal of frequency band 26 of subgroup 28 to the second decomposition filter group 24. In this case, Figure 4 The function of hearing aid 2 in the middle corresponds to that in Figure 2 The implementation shown is different. Instead, when control signal S2 is not applied to signal switcher 66 (and therefore the prediction is not accurate enough), signal switcher 66 directly guides the sub-band signals of frequency band 26 of subgroup 28 to sub-unit 36 ​​of data processing unit 34 for a given time period. In this case, all sub-band signals of the first frequency band 26 are processed, and in particular amplified in a frequency-specific manner, without further frequency splitting. Thus, no prediction is performed. After a predetermined time period, the second frequency split is activated, and therefore prediction is also activated to recheck the accuracy of the prediction by means of evaluation module 68.

[0088] In the previously described implementation variant, no voice recognition module 64 is included. Instead, the signal switcher 66 is controlled only via control signal S2.

[0089] According to Figure 4 In another embodiment of the hearing aid 2, not only a speech recognition module 64 is provided, but also an evaluation module 68 is provided. Here, the signal switcher 66 is controlled not only by control signal S1 but also by control signal S2. Preferably, control signals S1 and S2 are combined in an AND operation, so that the second frequency split is activated by the signal switcher 66 only when not only control signal S1 is applied but also control signal S2 is applied, that is, when not only is spoken speech recognized in the input audio signal I, but the prediction has sufficient accuracy.

[0090] According to Figure 4 In another variation of the hearing aid 2 (not shown in detail), a characteristic parameter Q is calculated individually for each frequency band 26 of the first subgroup 28, specifically by determining the corresponding band-specific predictor gain and comparing it with a corresponding band-specific threshold. Here, the evaluation module 68 outputs the control signal S2 in a band-specific manner only for one or more frequency bands 26 for which the band-specific predictor gain exceeds the corresponding associated threshold. Correspondingly, the signal switcher 66 selectively activates the second frequency split only for the relevant band 26 or the relevant bands 26. In this variation of the hearing aid, the predictor 46 is preferably arranged between the signal switcher 66 and the second decomposition filter group 24.

[0091] According to Figure 4In another variation of the hearing aid 2, not shown in detail, when the speech recognition module 64 identifies the presence of spoken speech in the corresponding frequency band 26, it generates a control signal S1 in a frequency band-specific manner. In this case, the signal switcher 66 also correspondingly and selectively activates the second frequency split only for the relevant frequency band 26 or multiple relevant frequency bands 26.

[0092] According to Figure 4 In another variation of the hearing aid 2, not shown in detail, when the value of the characteristic parameter Q is below a threshold, the predictor 46 is only interrupted in the signal path connecting the microphone 6 and the earpiece 8, but the predictor 46 continues to operate in the background of signal processing (here, the prediction will have no effect on the output audio signal O). This is achieved, for example, by the signal switcher 66 adjusting according to... Figure 4 The diagram is connected between the second synthesis filter bank 42 and the first synthesis filter bank 44 in a mirror-symmetric manner. In this case, the predictor 46 also outputs the feature parameter Q continuously when the feature parameter Q does not exceed the threshold. Here, it is not necessary, and therefore not set according to... Figure 4 The embodiment describes reactivating the signal switch 66 after a pre-given time period in order to recheck the accuracy of the prediction.

[0093] Preferably, it is implemented as software executed in the signal processor 12 during the operation of the hearing aid 2. Figures 2 to 4 The signal processor 12 shown includes components such as a decomposition filter bank device 20 having decomposition filter banks 22, 24, and possibly 50; a data processing unit 34 having subunits 36, 38, and possibly 58; a synthesis filter bank device 40 having synthesis filter banks 42, 44, and possibly 60; a predictor 46 and possibly a predictor 62; a possible speech recognition module 64; a signal switcher 66; and an evaluation module 68. Alternatively, one or more of these components may also be formed by non-programmable electronic circuitry.

[0094] The invention will become particularly apparent from the embodiments described above, but the invention is not limited to these embodiments. Rather, other embodiments of the invention can be derived from the foregoing description. In particular, individual features of the invention described by way of embodiments can also be combined in other ways without departing from the subject matter of the invention.

[0095] List of reference numerals

[0096] 2 hearing aids

[0097] 4 housings

[0098] 6 microphones

[0099] 8 earpieces

[0100] 10 batteries

[0101] 12 signal processors

[0102] 14 audio channels

[0103] 16 tips

[0104] 20-Decomposition Filter Bank Device

[0105] 22 (First) Decomposition Filter Bank

[0106] 24 (Second) Decomposition of Filter Bank

[0107] 26 (First) Band

[0108] 28 (First) Subgroup

[0109] 30 (Second) Band

[0110] 32 (Second) Subgroup

[0111] 34 Data Processing Units

[0112] 36 subunits

[0113] 38 subunits

[0114] 40 Synthetic Filter Bank Device

[0115] 42 (Second) Synthetic Filter Bank

[0116] 44 (First) Synthetic Filter Bank

[0117] 46 predictors

[0118] 46' Predictor (Replacement Position)

[0119] 50 (Third) Decomposition Filter Bank

[0120] 52 (First) Subgroup

[0121] 54 (Third) Band

[0122] 56 (Second) Subgroup

[0123] 58 subunits

[0124] 60 (Third) Synthetic Filter Bank

[0125] 62 predictors

[0126] 64 speech recognition module

[0127] 66 signal switcher

[0128] 68 assessment modules

[0129] I input audio signal

[0130] F error

[0131] O outputs audio signal

[0132] S1 control signal

[0133] S2 control signal

[0134] U supply voltage

Claims

1. A method for processing an input audio signal (I), - wherein, in a first frequency splitting, the input audio signal (I) is divided into a plurality of first frequency bands (26) by means of a first analysis filter bank (22), - wherein, in at least one further frequency splitting, first frequency bands (26) of a first subset (28) of the first frequency bands (26) are divided into a plurality of subbands (30, 54) by means of at least one further analysis filter bank (24), - wherein the input audio signal (I) divided into first frequency bands (26) or subbands (30, 54) is processed in a frequency-selective manner, and - wherein the input audio signal (I) divided into first frequency bands (26) or subbands (30, 54) and processed in a frequency-selective manner is combined into an output audio signal (O), characterized in that a prediction is applied to the first frequency bands (26) of the first subset (28) and / or to the subbands (30, 54) derived therefrom, in order to compensate for a delay difference between the first frequency bands (26) and the subbands (30) resulting from the further frequency splitting or each further frequency splitting.

2. The method according to claim 1, wherein the method for processing an input audio signal (I) is performed in a hearing device (4).

3. The method according to claim 1, wherein, the input audio signal (I) divided into first frequency bands (26) or subbands (30, 54) is amplified in a frequency-selective manner.

4. The method according to claim 1, wherein the first subset (28) of first frequency bands (26) is formed by a plurality of first frequency bands (26) whose respective center frequencies are directly adjacent and which include the lowest first frequency band (26).

5. The method according to claim 1, wherein a second subset (32) of first frequency bands (26) is processed in a frequency- selective manner without further frequency splitting.

6. The method according to claim 1, wherein the first frequency bands (26) have a uniform first bandwidth.

7. The method according to claim 1, wherein the prediction applied to the first frequency bands (26) of the first subset (28) or to the subbands (30, 54) is non-linear.

8. The method according to claim 1, wherein the prediction applied to the first frequency bands (26) of the first subset (28) or to the subbands (30, 54) is adaptive during signal processing.

9. The method according to any one of claims 1 to 8, wherein in at least one of the first frequency bands (26) or subbands (30, 54), the input audio signal (I) is analyzed as to whether voiced speech is present, wherein the further frequency splitting or each further frequency splitting is performed only if the presence of voiced speech is identified in the input audio signal (I).

10. The method according to any one of claims 1 to 8, wherein determining the accuracy of the prediction, and wherein the further frequency split or each further frequency split is performed only if the accuracy of the prediction meets a predefined criterion in at least one of the first frequency bands (26) or sub-bands (30, 54).

11. A device for processing an input audio signal (I), - having a first decomposition filter bank (22) configured for splitting the input audio signal (I) into a plurality of first frequency bands (26) in a first frequency split, - having at least one further decomposition filter bank (24, 50) connected downstream of the first decomposition filter bank (22) and configured for splitting a first subset (28) of the first frequency bands (26) into a plurality of sub-bands (30, 54) in at least one further frequency split, - having a signal processing unit (34) for frequency-selectively processing the input audio signal (I) split into first frequency bands (26) or sub-bands (30, 54), and - having a synthesis decomposition filter arrangement (40) connected downstream of the signal processing unit (34) and configured for combining the input audio signal (I) split into first frequency bands (26) or sub-bands (30, 54) and processed frequency-selectively into an output audio signal (O), characterized in that at least one predictor (46, 62) configured for applying a prediction to the first frequency bands (26) of the first subset (28) and / or to the sub-bands (30, 54) derived therefrom for compensating for a delay difference between the first frequency bands (26) and the sub-bands (30, 54) resulting from the further frequency split or each further frequency split.

12. The device according to claim 11, wherein the device is a hearing device.

13. The device according to claim 11, the signal processing unit (34) is configured for frequency-selectively amplifying the input audio signal (I) split into first frequency bands (26) or sub-bands (30, 54).

14. The device according to claim 11, wherein the first subset (28) of first frequency bands (26) is formed by a plurality of first frequency bands (26) whose respective center frequencies are directly adjacent and include the lowest first frequency band (26).

15. The device according to claim 11, wherein a second subset (32) of first frequency bands (26) is directly fed to the signal processing unit (34) for processing the second subset (32) of first frequency bands (26) frequency-selectively without further frequency splitting.

16. The device according to claim 11, wherein the first frequency bands (26) have a uniform first bandwidth.

17. The device according to claim 11, wherein the at least one predictor (46, 62) is a non-linear predictor.

18. The device according to claim 11, wherein the at least one predictor (46, 62) is an adaptive predictor during signal processing.

19. The device according to any one of claims 11 to 18, having a speech recognition module (64) configured for analyzing the input audio signal (I) for the presence of voiced speech, and having a switching device (66) configured for activating the at least one further decomposition filter bank (24, 50) only when the speech recognition module (64) recognizes the presence of voiced speech in the input audio signal (I).

20. The device according to any one of claims 11 to 18, having a switching device (66) configured for activating or deactivating the at least one further decomposition filter bank (24, 50) depending on the accuracy of the prediction.

Citation Information

Patent Citations

  • Method for optimising a multi-stage filter bank and corresponding filter bank and hearing aid

    EP2124335B1

  • Method for reducing latency of a filter bank for filtering an audio signal and method for low latency operation of a hearing system

    EP3197181A1

  • Voice information processing method and device, equipment and medium

    CN111128174A

  • Method for optimizing a multilevel filter bank and corresponding filter bank and hearing apparatus

    US20090290737A1

  • Systems and methods for reconstructing decomposed audio signals

    US20100094643A1