Speech enhancement using active masking control
The speech intelligibility enhancement system improves speech clarity in noisy environments by using in-ear headphones to manage sound pressure levels and compensate for vowel frequencies, addressing the masking issue in existing technologies.
Patent Information
- Application Number
- JP2025524177
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-31
- Filing Date
- 2023-10-30
- Publication Date
- 2025-10-20
AI Technical Summary
Existing noise suppression algorithms and hearing aids fail to improve speech recognition and intelligibility in noisy environments due to processing artifacts, unnatural sounds, and the masking effect of vowel frequencies on consonant frequencies, leading to increased mental effort in understanding speech.
A speech intelligibility enhancement system with in-ear headphones that utilize an acoustic pathway and an electroacoustic path to transmit and reproduce vowel and consonant dominant frequency ranges, respectively, improving the signal-to-masking ratio by compensating for vowel contributions in the acoustic path.
Enhances speech intelligibility by reducing the masking effect of vowels on consonants, allowing for clearer understanding in noisy environments through differential sound pressure level management.
Smart Images

Figure 2025534895000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a speech intelligibility enhancement system for difficult acoustic conditions and a method for enhancing speech intelligibility in difficult acoustic conditions. [Background technology]
[0002] Difficulty with voice communication in noisy environments is a common experience. Cocktail parties, cafes, and similar situations pose particular challenges because the signals (the voices of interlocutors) are very similar and often quieter than the noise (the multiple crosstalk of other people). For people with normal hearing, distinguishing between words requires a great deal of mental effort, and for people with very mild hearing loss, even more.
[0003] Many noise suppression algorithms (including adaptive microphone directional patterns) show substantial gains in signal-to-noise ratio (SNR), however, they often fail to provide better speech recognition scores in real tests due to, for example, processing artifacts and unnatural sounds.
[0004] Traditional passive hearing protectors generally attenuate too much, especially at higher frequencies, further worsening speech recognition. Additionally, traditional hearing protectors cause occlusion (i.e., users perceive their own voice as "hollow" or "booming" due to blocking the ear canal without compensation).
[0005] So-called musician's earplugs, which aim to attenuate a wide range of audio frequencies relatively equally so as not to distort the perception of music, also generally provide too much attenuation to be useful for understanding speech in noisy environments. They also often do not address the occlusion effect.
[0006] Hearing aids, on the other hand, aim to improve audibility by using general sound amplification techniques. This, as explained above, may not be helpful for people with normal or near-normal hearing who have difficulty understanding speech in noisy environments. To allow users to participate in conversations, most hearing aids incorporate a vent to allow bone / tissue-conducted sound from the user's own voice to escape the ear canal. However, this has the inherent problem that when the vent is large enough to provide acceptable perception of the user's own voice, much low-frequency energy from the surroundings enters the ear, is amplified by Helmholtz resonance, and masks important higher-frequency audio cues. To counteract this masking effect, high-frequency gain must be increased. This means that the overall level at the eardrum increases above the level that would have resulted in an open ear. However, when the level increases above a certain level (corresponding to approximately 65 dBA outside the ear for a subject with normal hearing) that is well below the level present at a typical party, frequency discrimination and speech understanding deteriorate.
[0007] For people with normal or near-normal hearing, an ear device that addresses one or more of the above-mentioned problems to improve listening comfort and / or speech recognition in noisy environments would be highly advantageous and useful. Summary of the Invention
[0008] The inventors have identified the above mentioned problems and challenges relating to speech listening comfort and intelligibility, especially in noisy environments, and have subsequently made the invention described below.
[0009] An aspect of the present invention is a speech intelligibility enhancement system for difficult acoustic conditions, the speech intelligibility enhancement system comprising at least one in-ear headphone device for insertion into a person's ear canal, the at least one in-ear headphone device being arranged with a portion facing the ear canal and a portion facing an environment, the at least one in-ear headphone device comprising: an acoustic pathway comprising a vent, the acoustic pathway connecting a portion facing the environment with a portion facing the ear canal; an electroacoustic path comprising a microphone in a portion facing the environment, a filter, and a loudspeaker in a portion facing the ear canal; the acoustic path is arranged to transmit acoustic sounds within a vowel dominant frequency range, and the electrical acoustic path is arranged to acoustically reproduce sound signals within a consonant dominant frequency range and within the vowel dominant frequency range; The electroacoustic path relates to a speech intelligibility enhancement system, wherein the electroacoustic path is arranged such that the signal-to-masking ratio is improved by the electroacoustic path compensating for contributions from the acoustic path in the vowel dominant frequency range.
[0010] Thereby, an advantageous system for enhancing speech intelligibility in difficult acoustic environments is provided, the benefits of which will become apparent below.
[0011] In this context, speech is understood as oral communication using a language, including non-tonal languages. Each language uses a combination of vowel and consonant sounds to form the sounds of its words. Vowels tend to be lower in frequency and louder than consonants and therefore carry most of the sound energy attributable to speech. However, in reality, it is the lower-energy, higher-frequency consonants that carry most of the meaning of words. Therefore, speech intelligibility is highly dependent on the frequency range of the sounds attributable to consonants. Consonants are more sensitive to upward masking than vowels, and therefore energy from vowels can impose a masking effect on consonants. Such masking effects are easily relatable, as they can occur when listening to someone speaking in a loud acoustic environment, such as a cafe with high levels of background noise. To overcome high levels of background noise, people tend to speak louder and with greater effort to be heard, a phenomenon often referred to as the Lombard effect. Speaking loudly has a significant impact on intelligibility for others because the added acoustic energy is concentrated around vowels, i.e., in the vowel-dominated frequency range, but adds very little energy to consonants, i.e., in the consonant-dominated frequency range. Thus, in social environments, everyone speaks louder, which means that a lot of energy is added to the vowel-dominated frequency range. This basically means that consonants are masked by vowels, resulting in difficulty in understanding what is being said. However, the importance of vowels in speech should not be underestimated; they still play a key role in speech intelligibility.
[0012] In this context, the signal-to-masking ratio is understood as a measure comparing the level of a desired signal with the level of a masking signal. In this context, the desired signal is a signal that is substantially within the consonant-dominant frequency range, and the masking signal is a signal that is substantially within the vowel-dominant frequency range. In other words, the signal-to-masking ratio can also be referred to as the consonant-to-vowel ratio. The masking signal does not necessarily represent an undesired sound, as is typical for noise signals when considering signal-to-noise ratios; however, the masking signal may actually contain audio cues that are beneficial to speech intelligibility. For example, the masking signal may include audio contributions from a speaker of interest (a person speaking to a person wearing a speech intelligibility enhancement system) and audio contributions made by multiple other people in the same acoustic environment as the speaker of interest and the person wearing the speech intelligibility enhancement system (this audio contribution may be referred to as multiple crosstalk noise throughout the following disclosure). In essence, the masking signal has a masking effect on consonants in the consonant-dominant frequency range, and by improving the signal-to-masking ratio, the masking effect can be reduced and speech intelligibility can be improved. It should be noted, therefore, that the speech intelligibility enhancement system is thereby effectively arranged to perform active masking control.
[0013] It should also be noted that the above discussion regarding improving the signal-to-masking ratio should be interpreted as improving in the context of a situation where the user is not wearing a speech intelligibility enhancement system, i.e., where at least one in-ear headphone device (such as two in-ear headphone devices) is not inserted in the user's ear canal.
[0014] According to an embodiment, improving the signal-to-masking ratio comprises increasing the resulting sound pressure level present in the consonant dominant frequency range relative to the resulting sound pressure level present in the vowel dominant frequency range.
[0015] Improving the signal-to-masking ratio, or consonant-to-vowel ratio, may involve increasing the resulting sound pressure level in the consonant-dominant frequency range relative to the resulting sound pressure level in the vowel-dominant frequency range. This may involve amplifying the acoustic sounds present in the consonant-dominant frequency range, i.e., the electrical acoustic path is arranged to perform sound amplification in the consonant-dominant frequency range. However, this should not be interpreted as meaning that the improvement in the signal-to-masking ratio is achieved solely by adjusting the gain of the electrical acoustic path in the consonant-dominant frequency range, since the electrical acoustic path is still arranged to compensate for the contribution by the acoustic path in the vowel-dominant frequency range. Increasing the resulting sound pressure level present in the consonant-dominant frequency range relative to the resulting sound pressure level present in the vowel-dominant frequency range is advantageous in that the influence of masking signals on the signals that are most interesting to speech intelligibility, i.e., the signals in the consonant-dominant frequency range, may be reduced, thereby improving speech intelligibility.
[0016] According to an embodiment, said improving the signal to masking ratio comprises reducing the difference between the resulting sound pressure level in said ear canal contributed by said vowel dominant frequency range and the resulting sound pressure level in said ear canal contributed by said consonant dominant frequency range by using said electrical acoustic path to compensate for contributions from said acoustic path in said vowel dominant frequency range.
[0017] The speech intelligibility enhancement system according to the present embodiment is advantageous in that it reduces the difference between the sound pressure level (SPL) contributed by the vowel-dominant frequency range and the SPL contributed by the consonant-dominant frequency range. The sound pressure level is the most commonly used measure of sound wave intensity and is typically measured in decibels (dB). The reduction of the sound pressure level difference is performed by compensating for the contribution from the acoustic path using the electrical acoustic path. The purpose of the compensation is not to completely cancel the contribution from the acoustic path, because the vowel content of the sound contributed by the acoustic path remains important in the reproduction of sound in the ear canal of a person wearing an in-ear headphone device. Without the vowel content contributed by the acoustic path, the sound as experienced by a wearer of at least one in-ear headphone device would sound unnatural and lack important features. However, the purpose of the compensation is to reduce the influence of high-energy vowels relative to the influence of lower-energy consonants. Thereby, the consonant-dominant portions of speech may be promoted relative to the vowel-dominant portions of speech, thus improving the intelligibility of speech in many acoustic environments.
[0018] In other words, the effect of the compensation is that the total transfer function from the external acoustic environment to the ear canal, resulting from contributions from both the acoustic and electro-acoustic paths, exhibits a smaller difference between the sound pressure levels in the vowel dominant frequency range and the sound pressure levels in the consonant dominant frequency range compared to the difference between them when no compensation is applied.
[0019] According to an embodiment of the present invention, the reduction of the difference is obtained by compensating for the contribution from the acoustic path using a filter implemented in a signal processor. The signal processor may apply filtering to the signal recorded by the microphone, thereby providing a filtered signal for reproduction using a loudspeaker. The effect of the sound reproduction of the filtered signal is that the effect of the acoustic sounds contributed by the acoustic path is attenuated in a sub-range or in the full range of the vowel dominant frequency range.
[0020] In this context, a vowel-dominant frequency range is understood as a range of frequencies that are substantially dominated by the presence of frequency components that form part of a vowel. Furthermore, in this context, a consonant-dominant frequency range is understood as a range of frequencies that are dominated by the presence of frequencies that form part of a consonant. Those skilled in the art will readily understand that because any tone produced by a human can contain multiple harmonics, including a first harmonic (or fundamental) and a second, third, fourth, etc. (or overtones), no clear line can be drawn between vowel and consonant frequency components, since overtones of vowel frequency components may exist in higher frequency ranges, such as the consonant-dominant frequency range. However, those skilled in the art of phonetics will understand that consonant frequency components typically exist in frequencies (e.g., 2 kHz to 4 kHz) that exceed the frequency components that make up vowels, which exist in lower frequencies (e.g., within the range of 50 Hz to 1 kHz).
[0021] At least one in-ear headphone device, such as two in-ear headphone devices, of a speech intelligibility enhancement system (or hereinafter "system") is arranged to be inserted into a person's ear canal. When inserted into the ear canal, the in-ear headphone device has a portion facing toward the ear canal, i.e., an "ear canal-facing portion," and a portion facing toward the person's surrounding environment, i.e., an "ambient-facing portion." These two portions of the in-ear headphone device are connected by the presence of an acoustic path. An acoustic path is understood as a path through which acoustic sound can propagate. The acoustic path comprises a vent, which is a channel or duct with a specific geometric shape that can be determined by acoustic concerns. The vent effectively connects the portion facing the environment with the portion facing the ear canal, ensuring that acoustic sounds present in the environment can propagate into the person's ear canal. The acoustic path is arranged so that acoustic sounds of a specific range of frequencies can propagate through the acoustic path, while acoustic sounds of other frequencies can be blocked. These acoustic characteristics can be attributed to the geometry (shape, cross-sectional area, length) of the vent. Typically, in prior art systems, such as in-ear headphone devices for listening to music, such vents are used to reduce the impact of occlusion effects on the listening experience. However, as will become apparent below, the presence of the vent serves another purpose: the acoustic reproduction of sound in the person's ear canal. Nevertheless, a beneficial effect of the presence of the vent is that the acoustic path may reduce the impact of occlusion effects on the wearer's experience of their own voice.
[0022] In addition to having an acoustic path with a vent, the in-ear headphone device also includes an electro-acoustic path with a microphone in the portion facing the environment, a filter, and a loudspeaker in the portion facing the ear canal, which allows sound from the external acoustic environment to be electronically, e.g., digitally, processed and reproduced within the ear canal.
[0023] The speech intelligibility enhancement system may further advantageously provide multi-band (e.g., two-band) dynamic range compression, thereby facilitating differential compression within vowel-dominant and consonant-dominant ranges.
[0024] The speech intelligibility enhancement system is further advantageous in that it may provide, at least to some extent, a natural reproduction of acoustic sounds in the external environment, provided at least in part by an acoustic path that facilitates a natural reproduction in the ear canal of sounds present in the external acoustic environment.
[0025] An in-ear headphone device may be understood as a headphone device arranged to be worn by a user by fitting the device in the user's outer ear, such as the concha, next to the ear canal. An in-ear headphone device may further extend at least partially into the user's ear canal. An in-ear headphone device may typically be shaped to fit at least partially within the outer ear and / or ear canal, thereby ensuring a secure fit of the device in the user's ear. An in-ear headphone device may also be understood as an in-ear headphone, earplug, ear-canal headphone, plug-in earphone, or hearable.
[0026] According to an embodiment, the consonant dominant frequency range includes frequencies above the vowel dominant frequency range.
[0027] A consonant dominant frequency range may include frequencies above the vowel dominant frequency range. Whether or not there is overlap between the consonant dominant frequency range and the vowel dominant frequency range, the consonant dominant frequency range may still include frequencies that are not in the vowel dominant frequency range, and these frequencies are above the vowel dominant frequency range. From this, it is clear that the consonant dominant frequency range is a frequency range related to frequencies above the vowel dominant frequency range.
[0028] According to an embodiment, the acoustic path is arranged with an acoustic transfer function having a low-pass characteristic with a passband and a cutoff frequency, the vowel dominant frequency range including frequencies below the cutoff frequency and the consonant dominant frequency range including frequencies above the cutoff frequency.
[0029] The acoustic path may be configured with an acoustic transfer function having a low-pass characteristic with a passband and a cutoff frequency, where the vowel-dominant frequency range includes frequencies below the cutoff frequency and the consonant-dominant frequency range includes frequencies above the cutoff frequency. By implementing such a low-pass characteristic in its transfer function, the acoustic path is configured to pass acoustic sounds having frequencies below the cutoff frequency (in the passband) and reach the ear canal of a user wearing at least one in-ear headphone device. However, acoustic sounds having frequencies above the cutoff frequency are severely restricted when passing through the acoustic path. Those skilled in the art of acoustics will readily understand that acoustic sounds with frequencies above the cutoff frequency may pass along the acoustic path, but will be significantly hindered. Typically, the cutoff frequency is the frequency at which 3 dB of attenuation occurs. The vowel dominant frequency range includes frequencies below the cutoff frequency and the consonant dominant frequency range includes frequencies above the cutoff frequency; however, this does not exclude the possibility that the vowel dominant frequency range also includes frequencies above the cutoff frequency and the consonant dominant frequency range includes frequencies below the cutoff frequency.
[0030] According to an embodiment, the cutoff frequency is in the range of 250 Hz to 4 kHz.
[0031] The cut-off frequency may be within the range of 250 Hz (Hertz) to 4 kHz (Kilohertz), such as 500 Hz to 2 kHz, such as 650 Hz to 1600 Hz, such as 700 Hz to 1200 Hz, for example 800 Hz, 900 Hz or 1 kHz.
[0032] According to an embodiment, the vowel dominant frequency range includes frequencies within the range of 50 Hz to 1 kHz.
[0033] The vowel dominant frequency range may include frequencies within the range of 50 Hz to 1 kHz, frequencies within the range of 400 Hz to 800 Hz, for example 600 Hz.
[0034] According to an embodiment, the consonant dominant frequency range includes frequencies within the range of 2 kHz to 4 kHz.
[0035] According to an embodiment, the difference is less than 15 dB, such as less than 10 dB, such as less than 8 dB, such as less than 6 dB, for example less than 5 dB.
[0036] The difference between the sound pressure level in the ear canal contributed by the vowel dominant frequency range and the resulting sound pressure level in the ear canal contributed by the consonant dominant frequency range may, after compensation, be less than 15 dB (decibels), such as less than 10 dB, such as less than 8 dB, such as less than 6 dB, for example less than 5 dB.
[0037] Reducing the difference between the sound pressure level contributed by the vowel dominant frequency range and the sound pressure level contributed by the consonant dominant frequency range so that the difference is less than 15 dB is considered a mere reduction of the impact provided by the acoustic path on the listening experience, and not a complete elimination of the impact provided by the acoustic path. In the present context, acoustic sounds propagating through the acoustic path should not be considered as unwanted sounds; quite the contrary, acoustic sounds propagating through the acoustic path from the environment-facing part of the in-ear headphone device to the ear canal-facing part of the in-ear headphone device may contribute to speech intelligibility.
[0038] According to an embodiment, the electroacoustic path is arranged to compensate for contributions from the acoustic path within the signal processing frequency range of 300 Hz to 1 kHz.
[0039] The electrical acoustic path may be arranged to compensate for contributions from that acoustic path within a signal processing frequency range using a signal processor, which may be a frequency range of 300 Hz to 1 kHz, a frequency range of 400 Hz to 800 Hz, for example 600 Hz.
[0040] According to an embodiment, the compensation of the contribution from the acoustic path is signal dependent.
[0041] By signal-dependence is understood at least that the signal processing, i.e., compensation, is dependent on the acoustic signals present in the external acoustic environment. Such signal-dependent compensation is advantageous in that the speech intelligibility enhancement system can better adapt to the external acoustic environment, thereby providing an improved listening experience.
[0042] According to an embodiment, the compensation of the contribution from the acoustic path is level dependent.
[0043] Level dependency is understood to mean the dependence of signal processing, i.e., compensation, on the sound pressure level measured, for example, in a vowel-dominant frequency range, e.g., the center frequency of the vowel-dominant frequency range. Such dependence is advantageous in that the speech intelligibility enhancement system can better adapt to the sound pressure level present in the external acoustic environment. For example, when the sound pressure level in the external acoustic environment is low, the compensation requirement to achieve a high level of speech intelligibility may be lower than when the sound pressure level is high. At such low sound pressure levels, the compensation may be kept at a low level, thereby not being seriously affected by the acoustic sound contributed by the acoustic path, for example, through less distortion. This achieves the advantage that the quality of the acoustic sound reproduction is always as high as possible in the ear canal of a user wearing at least one in-ear headphone device.
[0044] It should further be noted that a device implementing any of the above provisions may be arranged to intermittently adapt the compensation according to changes in the external acoustic environment, including adjusting the gain of the transfer function and even switching the compensation on and off. The device may additionally be arranged to perform other types of audio processing according to other acoustic conditions. Such other types of processing may include low-frequency amplification, which may be advantageous in quiet conversations.
[0045] According to an embodiment, the electroacoustic path is arranged to compensate for the contribution from the acoustic path by reproducing sound signals in at least part of the vowel dominant frequency range.
[0046] The electrical acoustic path may be arranged to reproduce sound signals of the external acoustic environment in at least a portion of the vowel-dominant frequency range. This may include, for example, reproducing the sound signals using loudspeakers of an in-ear headphone device, such that compensation for the contributions by the acoustic path is achieved. The compensation may include reproducing sound signals having an opposite polarity or a different phase than the audio signals contributed by the acoustic path in at least the vowel-dominant frequency range.
[0047] According to an embodiment, the electrical acoustic path is arranged to reproduce sound signals in the at least part of the vowel dominant frequency range having a polarity opposite to that of the acoustic sounds transmitted by the acoustic path, the effect of compensation being that the perceived loudness of acoustic sounds in the vowel dominant frequency range is reduced compared to a situation in which at least one in-ear headphone device is not inserted in the ear canal of the user / wearer.
[0048] According to an embodiment, the electrical acoustic path is arranged to reproduce sound signals in said at least part of said vowel dominant frequency range by applying a phase shift to the sound signals.
[0049] Since the electrical acoustic path comprises a microphone, or multiple microphones according to other embodiments, and a loudspeaker, the electrical acoustic path may perform signal processing on the recorded signals. The signal processing may include applying a phase shift to such signals. The electrical acoustic path may thereby be arranged to reproduce, in the user's ear canal, sound signals originating from the external acoustic environment in at least a portion of a vowel-dominant frequency range by applying a phase shift. The application of a phase shift may result in a reproduced signal having a canceling effect on audio signals transmitted from the external acoustic environment to the ear canal via the acoustic path and the vent.
[0050] According to an embodiment, the phase shift is greater than 90 degrees and less than 270 degrees.
[0051] The applied phase shift may be greater than 90 degrees and less than 270 degrees. The phase shift may be applied to any frequency within the vowel dominant frequency range, such as the center frequency of the vowel dominant frequency range, for example, a frequency of 600 Hz.
[0052] According to an embodiment of the present invention, the microphone and the speaker are oppositely wired with respect to the positive and negative terminals.
[0053] According to an embodiment, the vent is a damped vent.
[0054] The vent may be a damped vent comprising one or more vent elements and one or more damping elements, such as a vent having a damping fabric located at one or both ends of the vent, or a vent configured with integrated damping.
[0055] The addition of a non-damped vent to an in-ear headphone device may reduce the occlusion effect when the in-ear headphone device is worn by a user, but results in Helmholtz resonance. By further adding a damping element to the vent, thereby providing a damped vent, Helmholtz resonance and the associated distortion it may generate can be eliminated.
[0056] According to an embodiment, the loudspeaker and the vent are acoustically separated within the at least one in-ear headphone device.
[0057] According to an embodiment, the vent is disposed with a cross-sectional area equivalent to a cylinder having a diameter in the range of 1.5 mm to 3.5 mm, such as 2.0 mm to 3.0 mm, for example 2.3 mm or 2.5 mm.
[0058] A preferred cross-sectional area for the vent is, for example, 1.8 mm 2 (square millimeter) ~ 9.6mm 2 , 3.1mm 2 ~7.1mm 2 For example, 4.2 mm 2 or 4.9 mm 2 The vent may have a variety of cross-sectional shapes, such as circular, rectangular, and semicircular, may have a cross-sectional area that varies along its length, or may be combined by two or more vents or split vents, but is preferably designed with dimensions comparable to those described above for a cylindrical vent.
[0059] According to an embodiment, the vent is arranged with a length equivalent to a cylinder having a length in the range of 2.5 mm to 10 mm, such as 3.5 mm to 9 mm, such as 4.5 mm to 8 mm, for example 5 mm or 7 mm.
[0060] The vent may have a variety of shapes along its length, may be straight, curved, or bent, may be combined by two or more vents or split vents, but is preferably designed with dimensions comparable to those described above for a cylindrical vent.
[0061] According to an embodiment, the filter is arranged in a signal processor, such as a digital signal processor, of the at least one in-ear headphone device.
[0062] According to an embodiment, the at least one in-ear headphone device is battery powered, such as powered by a rechargeable battery.
[0063] According to an embodiment, the at least one in-ear headphone device comprises two in-ear headphone devices, one for each ear canal of the person, the two in-ear headphone devices being arranged to coordinate settings between them.
[0064] The speech intelligibility enhancement system may include two in-ear headphone devices, one for each ear canal of a person, with the two devices arranged to coordinate settings between them. Thereby, a speech intelligibility enhancement system having the same advantages as described above and suitable for simultaneous use with both ears of a user is achieved. It should be noted that any effects and advantages described with respect to at least one in-ear headphone device apply equally to both in-ear headphone devices of this embodiment.
[0065] According to an embodiment, the at least one in-ear headphone device comprises a feedback microphone in the part facing the ear canal.
[0066] At least one in-ear headphone device of the voice intelligibility enhancement system may comprise a feedback microphone disposed in a portion of the at least one in-ear headphone device facing the ear canal. The feedback microphone is advantageous in that it facilitates improved control of the voice processing performed by the electro-acoustic path of the at least one in-ear headphone device. In particular, the feedback microphone may be used to adapt the feedforward processing of the electro-acoustic path. The feedback is further advantageous in that it enables the voice intelligibility enhancement system to detect whether a user / wearer of the system is speaking and to adapt the electro-acoustic path accordingly to provide the user / wearer with a desired impression of the wearer's own voice.
[0067] According to an embodiment, the electroacoustic path is arranged to compensate for contributions from the acoustic path based on input provided by the feedback microphone.
[0068] Compensating for contributions from the acoustic path based on input provided by the feedback microphone is advantageous in that it improves control of the audio processing performed by the electro-acoustic path of the at least one in-ear headphone device. In particular, basing the compensation on input provided by the feedback microphone may ensure that the acoustic sound present in a user's ear canal when the at least one in-ear headphone device is inserted therein actually reflects a desired listening experience.
[0069] According to an embodiment, the microphone of the electroacoustic path is a directional microphone.
[0070] In a preferred embodiment of the present invention, the microphone in the electrical acoustic path of at least one in-ear headphone device is a directional microphone. A directional microphone is understood to be a microphone that is most sensitive in one or more directions. In other words, a directional microphone has a polar pattern other than omnidirectional. Those skilled in the art will readily understand that such a directional microphone can be realized in numerous ways, including the use of multiple microphones arranged in a specific configuration, or by using a single microphone in conjunction with multiple microphone ports / ducts. When implemented in at least one in-ear headphone device, a directional microphone is advantageous because omnidirectional sound contributions, such as multiple crosstalk noise, can be suppressed relative to sound contributions with more directional characteristics, such as the relevant speech of a speaker standing in front of the user / wearer of the speech intelligibility enhancement system. This can further improve speech intelligibility.
[0071] According to an embodiment of the present invention, the directional microphone has a sharp directional characteristic.
[0072] According to an embodiment, the electro-acoustic path of the at least one in-ear headphone device comprises a plurality of microphones.
[0073] According to embodiments of the present invention, the electrical acoustic path of the at least one in-ear headphone device may comprise a plurality of microphones, such as two or more microphones, which may be arranged such that the at least one in-ear headphone device comprises a directional microphone and an omnidirectional microphone.
[0074] According to an embodiment, the electroacoustic path is arranged to amplify sound with a nominal gain in the passband of the electroacoustic path.
[0075] The electroacoustic path may be arranged to amplify sound with a nominal gain in the passband of the electroacoustic path, such as to amplify sound with a nominal gain across the entire passband of the electroacoustic path, which is advantageous in situations of low sound pressure levels where speech understanding may be difficult.
[0076] Another aspect of the present invention is a method for enhancing speech intelligibility in difficult acoustic conditions, the method comprising: inserting at least one in-ear headphone device into the person's ear canal, the at least one in-ear headphone device being arranged with an ear canal-facing portion and an environment-facing portion, the at least one in-ear headphone device comprising an acoustic path comprising a vent coupling the environment-facing portion with the ear canal-facing portion, and an electro-acoustic path comprising a microphone in the environment-facing portion, a filter, and a loudspeaker in the ear canal-facing portion; transmitting acoustic sounds within a vowel dominant frequency range through the acoustic pathway from a portion facing the environment to a portion facing the ear canal; acoustically reproducing, via said electrical acoustic pathway, sound signals within a consonant dominant frequency range and within said vowel dominant frequency range; and compensating for contributions from said acoustic path in said vowel dominant frequency range by said electro-acoustic path so that the signal to masking ratio is improved.
[0077] Thereby, a method for enhancing speech intelligibility in difficult acoustic conditions is provided, which is advantageous for at least the same reasons as given with respect to the speech intelligibility enhancement system above.
[0078] According to an embodiment, the method is performed by a speech intelligibility enhancement device according to any of the preceding provisions. [Brief explanation of the drawings]
[0079] Various embodiments of the present invention are described below with reference to the drawings. [Figure 1] 1 illustrates an in-ear headphone device of a speech intelligibility enhancement system according to an embodiment of the present invention. [Figure 2a] 1 illustrates various in-ear headphone devices according to embodiments of the present invention. [Figure 2b] 1 illustrates various in-ear headphone devices according to embodiments of the present invention. [Figure 2c] 1 illustrates various in-ear headphone devices according to embodiments of the present invention. [Figure 2d] 1 illustrates various in-ear headphone devices according to embodiments of the present invention. [Figure 3a] 10A-10C illustrate various layouts of acoustic path vents suitable for use in in-ear headphone devices according to embodiments of the present invention. [Figure 3b] 10A-10C illustrate various layouts of acoustic path vents suitable for use in in-ear headphone devices according to embodiments of the present invention. [Figure 3c] 10A-10C illustrate various layouts of acoustic path vents suitable for use in in-ear headphone devices according to embodiments of the present invention. [Figure 3d] 10A-10C illustrate various layouts of acoustic path vents suitable for use in in-ear headphone devices according to embodiments of the present invention. [Figure 3e] 10A-10C illustrate various layouts of acoustic path vents suitable for use in in-ear headphone devices according to embodiments of the present invention. [Figure 3f] 10A-10C illustrate various layouts of acoustic path vents suitable for use in in-ear headphone devices according to embodiments of the present invention. [Figure 3g] 10A-10C illustrate various layouts of acoustic path vents suitable for use in in-ear headphone devices according to embodiments of the present invention. [Figure 3h] 10A-10C illustrate various layouts of acoustic path vents suitable for use in in-ear headphone devices according to embodiments of the present invention. [Figure 4]1 illustrates characteristics of acoustic and electro-acoustic paths according to an embodiment of the present invention. [Figure 5] Illustrates the concept of using an electroacoustic path to compensate for contributions from acoustic paths in the vowel dominant frequency range. [Figure 6] 1 illustrates the spectrum of a sound signal present in a room, the transfer function of an in-ear headphone device according to an embodiment of the present invention, and the application of the transfer function to a sound signal that is useful for understanding the present invention. [Figure 7] 1 illustrates the spectrum of a sound signal present in a room, the transfer function of an in-ear headphone device according to an embodiment of the present invention, and the application of the transfer function to a sound signal that is useful for understanding the present invention. [Figure 8] 1 illustrates the spectrum of a sound signal present in a room, the transfer function of an in-ear headphone device according to an embodiment of the present invention, and the application of the transfer function to a sound signal that is useful for understanding the present invention. [Figure 9] 1 illustrates the spectrum of a sound signal present in a room, the transfer function of an in-ear headphone device according to an embodiment of the present invention, and the application of the transfer function to a sound signal that is useful for understanding the present invention. [Figure 10] 1 illustrates the spectrum of a sound signal present in a room, the transfer function of an in-ear headphone device according to an embodiment of the present invention, and the application of the transfer function to a sound signal that is useful for understanding the present invention. [Figure 11] 1 illustrates the spectrum of a sound signal present in a room, the transfer function of an in-ear headphone device according to an embodiment of the present invention, and the application of the transfer function to a sound signal that is useful for understanding the present invention. [Figure 12] 1 illustrates the spectrum of a sound signal present in a room, the transfer function of an in-ear headphone device according to an embodiment of the present invention, and the application of the transfer function to a sound signal that is useful for understanding the present invention. [Figure 13] 1 illustrates the spectrum of a sound signal present in a room, the transfer function of an in-ear headphone device according to an embodiment of the present invention, and the application of the transfer function to a sound signal that is useful for understanding the present invention. [Figure 14]1 illustrates the spectrum of a sound signal present in a room, the transfer function of an in-ear headphone device according to an embodiment of the present invention, and the application of the transfer function to a sound signal that is useful for understanding the present invention. [Figure 15] 1 illustrates the spectrum of a sound signal present in a room, the transfer function of an in-ear headphone device according to an embodiment of the present invention, and the application of the transfer function to a sound signal that is useful for understanding the present invention. [Figure 16] 1 illustrates the spectrum of a sound signal present in a room, the transfer function of an in-ear headphone device according to an embodiment of the present invention, and the application of the transfer function to a sound signal that is useful for understanding the present invention. [Figure 17] 1 illustrates a speech intelligibility enhancement system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0080] 1 illustrates a speech intelligibility enhancement system 101 according to an embodiment of the present invention. The speech intelligibility enhancement system 101 is shown as comprising an in-ear headphone device 102; however, according to another embodiment, the speech intelligibility enhancement system 101 may comprise two in-ear headphone devices 102, one for each ear of a person. Thus, the following description regarding the in-ear headphone device 102 applies equally to a system comprising two in-ear headphone devices.
[0081] 1 shows the in-ear headphone device 102 when inserted into the ear canal 109 of a person / user wearing the in-ear headphone device. The in-ear headphone device 102 preferably includes flexible ear tips 111 for resting on the user's outer ear 110 and providing acoustic sealing within the different user's ear canal 109.
[0082] The in-ear headphone device 102 comprises a microphone 103 arranged to primarily record acoustic sounds from an external acoustic environment 108. In the drawings of this embodiment, the microphone 103 is shown arranged at the end of the in-ear headphone device 102 facing the external acoustic environment; however, in other embodiments of the invention, the microphone 103 may be arranged further within the in-ear headphone device 102 and acoustically coupled to the external acoustic environment 108 by a microphone conductor (not shown). The in-ear headphone device further comprises a signal processor 104 in the form of a digital signal processor configured to receive the recorded audio signal from the microphone 103 and apply a filter thereto (a digital filter in this embodiment) to provide a filtered audio signal for sound reproduction using a loudspeaker 105 of the in-ear headphone device. In the drawings of this embodiment, a loudspeaker 105 is shown included within the in-ear headphone device 102, with acoustic sounds emitted by the loudspeaker 105 being transmitted to the ear canal 109 via a loudspeaker duct 106. However, the loudspeaker duct 106 may be omitted in other embodiments, and the loudspeaker 105 may be disposed closer to the end of the in-ear headphone device 102 that faces the ear canal. The assembly including the microphone 102, the signal processor 104, and the loudspeaker 105 is hereinafter referred to as the electro-acoustic path.
[0083] In addition to the electrical acoustic path, the in-ear headphone device 102 comprises an acoustic path comprising a vent 107. A vent is a narrow duct along which acoustic sound can propagate. The purpose of the vent is to facilitate the transmission of low-frequency acoustic sound between the ear canal 109 and the external acoustic environment 108. In other words, the vent 107 facilitates coupling between the portion of the in-ear headphone device 102 facing the environment and the portion of the in-ear headphone device 102 facing the ear canal. The boundary between the portion of the in-ear headphone device 102 facing the ear canal and the portion facing the environment is at the periphery of the in-ear headphone device 102, where it is generally in contact with the ear canal 109, i.e., it substantially blocks the ear canal.
[0084] 2a-2d illustrate various in-ear headphone devices 102 according to embodiments of the present invention.
[0085] 2a shows the in-ear headphone device 102 of FIG. 1 also inserted into a user's ear canal 109, according to an embodiment. As is clear from this figure, acoustic sounds present in the external acoustic environment 108 can propagate through an acoustic path in the in-ear headphone device 102, i.e., through the vent 107 and its vent element 202, into the user's ear canal 109. Furthermore, acoustic sounds present in the external acoustic environment 108 are picked up by the microphone 103, processed by the signal processor 104, and acoustically reproduced by the loudspeaker 105, from which the reproduced sounds are sent to the ear canal 109 via the loudspeaker duct 106. From this, it is clear that the overall transfer function of sound from the external acoustic environment 108 into the ear canal 109 comprises two contributions: an acoustic path and an electro-acoustic path. The sound picked up by the user's tympanic membrane 201 therefore arises from these contributions. As can be seen in this figure, the vent comprises a single vent element 202 in the form of a duct, however, other configurations of the vent are possible according to other embodiments, as will become apparent from the description below.
[0086] 2b shows a variation of the in-ear headphone device 102 as seen in FIG. 2a, according to another embodiment. In this embodiment, the vent 107 is a damped vent that additionally comprises a damping element 203. The damping element according to this embodiment is a damping fabric located at one end of the damped vent 107. In another embodiment, the damping properties of the damped vent 107 are provided by damping fabric at both ends of the damped vent 107, and in other embodiments, the damping properties of the damped vent 107 are provided by slits or openings in the vent element 202.
[0087] FIG. 2c shows yet another variation of the in-ear headphone device 102 as seen in FIG. 2a, according to another embodiment. In this embodiment, the in-ear headphone device includes a feedback microphone 204 in addition to the microphone 103. The feedback microphone is shown disposed immediately adjacent to the portion of the in-ear headphone device 102 that faces the ear canal; however, according to other embodiments, the feedback microphone 204 may be disposed further toward the center of the interior of the in-ear headphone device 102 and acoustically coupled to the ear canal 109 via a microphone conductor (not shown). The feedback microphone 204 is disposed to pick up sounds within the ear canal 109 and provide a recorded signal to the signal processor 104. Specifically, the feedback microphone may detect sound pressure levels throughout a frequency range, including at least low frequencies, such as frequencies in the range of 50 Hz to 1 kHz (an example of a vowel-dominant frequency range), and higher frequencies, such as frequencies in the range of 2 kHz to 4 kHz (an example of a consonant-dominant frequency range). Typically, such a microphone will be configured to detect at least the full frequency range audible to humans (ie, the audible range), which is typically frequencies within the range of 20 Hz to 20 kHz.
[0088] Figure 2d shows another embodiment that is a variation of the in-ear headphone device 102 as seen in Figure 2c. As can be seen, the in-ear headphone device 102 comprises an attenuated vent 107 that comprises a vent element 202 and a damping element 203, similar to the attenuated vent 107 described in relation to Figure 2b. In another embodiment, the damping properties of the attenuated vent 107 are provided by damping fabric on both ends of the attenuated vent 107, while in other embodiments, the damping properties of the attenuated vent 107 are provided by slits or openings in the vent element 202.
[0089] 3a-3h illustrate various layouts of acoustic path vents 107 suitable for use in in-ear headphone devices 102 according to embodiments of the present invention. It should be noted that throughout the figures, damped vents are illustrated; however, according to other embodiments of the present invention, all of the illustrated vents may be used without damping elements.
[0090] 3a shows a side view of a damped vent 107 according to an embodiment of the present invention. The damped vent 107 comprises a vent element 202 in the form of a cylinder and a damping element 203 in the form of a damping fabric. The vent element 202 is illustrated in this embodiment as a cylindrical element, although other geometric shapes are contemplated.
[0091] The damping element 203 in the form of a damping fabric is illustrated as being located at one end of the vent element 202, however, it may be positioned at any end of the vent element 202, and in another embodiment of the present invention, the damped vent 107 comprises a damping element 203 at both ends of the damped vent 107. The damping element 203 in this embodiment is positioned within the opening of the vent element 202, however, in another embodiment of the present invention, the damping element 203 may be positioned to cover the opening of the vent element 202.
[0092] 3b shows a side view of a damped vent 107 according to an embodiment of the present invention. Several vent elements 202 form a branched damped vent 107 that further comprises a damping element 203 in the form of a damping fabric. The damping element 203 in this embodiment is positioned within the opening of the vent element 202; however, in other embodiments of the present invention, the damping element 203 may be positioned to cover the opening of the vent element 202. Furthermore, in other embodiments of the present invention, the branched damped vent may comprise any number of damping elements 203, such as damping elements 203 that cover all of the openings of the vent element 202.
[0093] Figures 3c and 3d show two side views of an attenuated vent 107 according to an embodiment of the present invention. Figure 5c shows an attenuated vent 107 constructed with a loudspeaker duct 106 to which a loudspeaker 105 can be acoustically coupled. In this embodiment of the present invention, the loudspeaker duct 106 and the attenuated vent 107 form a cylindrical acoustic tube, i.e., each of the two has a semi-cylindrical geometry. In other embodiments of the present invention, the loudspeaker duct 106 and the attenuated vent 107 may form a combined acoustic tube having any geometry. In Figure 3c, a dashed line cc is shown representing plane c. In Figure 3e, a view of the embodiment from plane c is illustrated, showing the longitudinal geometry of the combined loudspeaker duct 106 and attenuated vent 107.
[0094] Figure 3d illustrates an embodiment of the present invention in which an in-ear headphone device 102 (not shown) includes two separate attenuated vents 107. Each attenuated vent 107 is similar to the attenuated vent 107 shown in connection with the embodiment of Figure 3a. Similarly, the attenuated vent 107 configuration of Figure 3d includes a vent element 202 and a damping element 203. The damping element 203 in this embodiment is a damping fabric that resides in the opening of the vent element 202, although other configurations of damping elements are contemplated.
[0095] 3f illustrates an embodiment of the invention in which the damping properties of the damped vent 107 are enhanced by a damping element 203 in the form of a slit. In another embodiment, the damping element 203 is integrated into the vent element 202, for example to disrupt air flow or promote air leakage.
[0096] 3g illustrates an embodiment of the present invention in which a microphone, for example, a feedback microphone 204, is arranged to primarily record sound from the attenuated vent 107. The microphone may therefore be considered to be acoustically coupled to a vent element 202 of the attenuated vent 107 in the in-ear headphone device 102. In other embodiments, the in-ear headphone device 102 comprises several vent elements 202, and the microphone and / or loudspeaker may be coupled to any of these vent elements 202, according to embodiments of the present invention. In the embodiment shown in FIG. 3g, the attenuated vent 107 has a single damping element 203 on one side. In such an embodiment, the microphone may therefore primarily record sound from the external environment or primarily record sound from the ear canal, depending on the precise positioning of the damping element 203 and the microphone.
[0097] 3h illustrates an embodiment of the present invention in which the loudspeaker duct 106 and the attenuated vent are partially coupled by a damping element 203. The attenuated vent 107 also comprises damping elements 203 on either end of the vent element 202. The loudspeaker duct 106 and the attenuated vent 107 may feature any type of partition according to embodiments of the present invention. The loudspeaker 105 may, for example, be acoustically coupled to the attenuated vent 107 in the in-ear headphone device 102, acoustically isolated from the attenuated vent 107 in the in-ear headphone device 102 (see, for example, FIG. 3c), or partially coupled to the attenuated vent 107 in the in-ear headphone device 102, as illustrated in FIG. 3h.
[0098] In the above-described embodiments of the present invention, various configurations of the attenuated vent 107 are shown. However, the present invention is not limited to any particular configuration, and various other embodiments are available to those skilled in the art. An attenuated vent configuration may be realized by any combination of the above-described embodiments, and thus an attenuated vent configuration may include one or more attenuated vents 107, each attenuated vent may include any number of vent elements 202 and damping elements 203, microphones and / or loudspeakers may be acoustically coupled to the vent elements or may have individual ducts, and the vents and ducts may have any geometric shape. Furthermore, as already mentioned, all of the vents shown in FIGS. 3a-3h may be used without the damping element 203 according to other embodiments of the present invention.
[0099] FIG. 4 illustrates the characteristics of acoustic path 501 and electrical acoustic path 502 according to an embodiment of the present invention. The diagram shows a horizontal axis representing frequency (f) in units of Hertz (Hz). As can be seen in the diagram, the frequency axis includes two frequency ranges: a vowel dominant frequency range VDF and a consonant dominant frequency range CDF. The vowel dominant frequency range VDF includes frequencies in the range of 50 Hz to 800 Hz, and the consonant dominant frequency range includes frequencies in the range of 2000 Hz (2 kHz) to 4000 Hz (4 kHz). Although the two frequency ranges are illustrated as two separate ranges, this does not exclude that signal content related to vowels may exist outside the vowel dominant frequency range VDF and that signal content related to consonants may exist outside the consonant dominant frequency range CDF. In the present context, the consonant dominant frequency range CDF includes frequencies above those included in the vowel dominant frequency range VDF. In situations where there is party noise or the like, most of the noise energy falls within the vowel dominant frequency range VDF. This diagram also illustrates the passbands of acoustic path 501 and electrical acoustic path 502 of in-ear headphone device 102. Acoustic path 501 comprises at least vent 107 (see, for example, FIG. 1 ), and electrical acoustic path 502 comprises at least microphone 103, signal processor 104, and loudspeaker 105. Acoustic path 501 can be any of the acoustic paths described above, and electrical acoustic path 502 can be any of the electrical acoustic paths described above. As can be seen, acoustic path 501 is focused on the vowel dominant frequency range VDF. Acoustic path 501 is effectively a low-pass filter if the vowel dominant frequency range VDF is within the passband of acoustic path 501. However, electrical acoustic path 502 processes a much broader frequency range than acoustic path 501, encompassing both the vowel dominant frequency range VDF and the consonant dominant frequency range CDF. Figure 4 also illustrates a vertical arrow extending from electrical acoustic path 502 to acoustic path 501 within the vowel dominant frequency range VDF. The arrow represents the compensation performed by electrical acoustic path 502. This compensation is best understood by considering Figure 5.
[0100] FIG. 5 illustrates the concept of using electrical acoustic path 502 to compensate for contributions from acoustic path 501 in the vowel-dominant frequency range VDF. Speech contains both vowels and consonants, and speech intelligibility is largely due to the accurate detection of consonants. However, in many situations, such as a cocktail party, the vowels from competing speakers, which carry most of the acoustic energy of speech, have a masking effect on the consonants of their conversational partners. In other words, signal content within the vowel-dominant frequency range VDF can impose a masking effect on signal content present in the consonant-dominant frequency range CDF. For this reason, the in-ear headphone device of the speech intelligibility enhancement system (see, for example, the in-ear headphone device of Figures 1 and 2a-2d) is arranged so that the difference 503 between the resultant sound pressure level in the ear canal 109 contributed by the vowel dominant frequency range VDF and the resultant sound pressure level in the ear canal 109 contributed by the consonant dominant frequency range CDF is reduced by compensating for the contribution from acoustic path 501 in the vowel dominant frequency range VDF using electrical acoustic path 502. Figure 5 illustrates the resultant sound pressure level (SPL) 506 contributed by the vowel dominant frequency range VDF and the resultant sound pressure level 507 contributed by the consonant dominant frequency range. As seen in this embodiment, the resulting sound pressure level 506 contributed by the vowel dominant frequency range resides at the center frequency 504 of the vowel dominant frequency range VDF, and the resulting sound pressure level 507 contributed by the consonant dominant frequency range resides at the center frequency 505 of the consonant dominant frequency range CDF. However, in other embodiments, the resulting sound pressure level may represent the average sound pressure level of the entire vowel dominant and consonant dominant frequency ranges, or the average sound pressure level of a subrange thereof. The difference 503 between the two resulting sound pressure levels can be seen in FIG. 5. In preferred embodiments, the difference 503 is maintained below 15 dB, and in even more preferred embodiments, the difference 503 is maintained below 10 dB. Such maintenance may require that the difference 503 be reduced, which is achieved by using electrical acoustic path 502 to compensate for the contribution from acoustic path 501 within the vowel dominant frequency range VDF. Such compensation may be achieved in multiple ways according to embodiments of the present invention.In this embodiment, the signal processor 104 of the speech intelligibility enhancement system 101 is essentially arranged to apply a phase shift to the signal recorded by the microphone 103, thereby acoustically reproducing a phase-shifted audio signal within the vowel-dominant frequency range VDF using the loudspeaker 105. Importantly, the phase-shifted audio signal has an opposite effect on the acoustic sounds in the ear canal 109 contributed by the acoustic path 501, and this effect ensures that the overall transfer function of sound from the external acoustic environment 108 to the ear canal 109 exhibits a characteristic where the difference 503 is below a predetermined level. More importantly, the compensation is not intended to completely counter the acoustic sounds in the ear canal 109 contributed by the acoustic path 501, as the aim remains to achieve a degree of natural reproduction of the acoustic sounds within the passband of the acoustic path 501. This is particularly important as vowels produced by a conversation partner are themselves speech cues and also establish a time window for when important consonants may appear. Clearly, reducing the difference 503 as detailed above is a way to improve the signal to masking ratio.
[0101] In a preferred embodiment, the signal processing algorithm adjusts the amount of attenuation applied to the vowel dominant frequency range VDF depending on the sound pressure level, so that when low frequency levels are low enough that consonant masking is unlikely to occur, the sound is perceived with a natural and / or desired spectral balance.
[0102] In another preferred embodiment, the signal processing algorithm is arranged to detect when the wearer of the in-ear headphone device is speaking and adjust the compensation to maintain a natural impression of the wearer's own voice.
[0103] Figures 6 to 15 illustrate spectra for an in-ear headphone device 102 according to an embodiment of the present invention as shown in Figure 16. In this embodiment, the in-ear headphone device 102 of the speech intelligibility enhancement system 101 comprises two microphones. Further details about this embodiment are provided in the text accompanying Figure 16.
[0104] 6 illustrates the spectra of three signals as they would appear without the baffle effect, i.e., as if the signals were recorded using an omnidirectional microphone located where the wearer of the speech intelligibility enhancement system 101 would be standing. The figure shows three signal curves S1, S2, and S3 plotted on a graph showing amplitude spectral density (ASD) in units of 20 microPascals per square root of Hertz [dB re 20 μPa / sqrt(Hz)] as a function of frequency in units of Hertz [Hz].
[0105] The signal curve S1 represents the long-term average spectrum of the noise present in a room with 30 people speaking. Throughout the following description, this will be referred to as multiplex crosstalk noise.
[0106] Signal curve S2 also represents the long-term average spectrum of a single speaker located approximately one meter away from the wearer of the speech intelligibility enhancement system 101. Any speech pauses made by the speaker have been excluded from the integral leading to signal curve S2.
[0107] Signal curve S3 represents the short-time spectrum of the consonant "t" spoken by a single speaker located one meter away from the wearer of the speech intelligibility enhancement system 101. As can be seen, the spectrum of the consonant "t" peaks at approximately 3 kHz, i.e., within the consonant dominant frequency range CDF. Note that the consonant "t" was chosen for demonstration purposes only; one skilled in the art would have knowledge of the spectra of other consonants that could easily be used instead to demonstrate the same principles as those described below.
[0108] The three signal curves S1, S2, and S3 together represent a typical cocktail party situation where speech intelligibility is challenged by the presence of multiple crosstalk noises that have a masking effect on consonants important to speech intelligibility. The three signal curves S1, S2, and S3 may hereinafter be considered as input signals to signal processing by the acoustic and electroacoustic paths of at least one in-ear headphone device. This signal processing is illustrated by the transfer functions in Figures 7-9.
[0109] FIG. 7 illustrates four simplified transfer functions T1 (squares), T2 (crosses), T3 (triangles), and T4 (circles). The transfer functions are simplified in the sense that they do not take into account ear canal resonance. The transfer functions are plotted on a graph showing real-ear gain (REG) in decibels (dB) as a function of frequency in Hz. Transfer function T1 is the transfer function for the vent 107 of the acoustic path. Hereinafter, this transfer function will be referred to as the vent transfer function. Transfer function T2 is the transfer function of an audio signal recorded by one of the microphones of the in-ear headphone device, which in this case functions as a pressure microphone, e.g., an omnidirectional microphone. Hereinafter, this transfer function will be referred to as the omnidirectional microphone transfer function. Transfer function T3 is the transfer function of an audio signal recorded by one or more microphones of the in-ear headphone device that functions as a directional microphone of the acute directivity type. Because directional microphones are most sensitive in a particular direction, they are less sensitive to acoustic sounds with more diffuse characteristics, such as multi-channel crosstalk noise. This is reflected by transfer function T4, which is the transfer function of a directional microphone when subjected to diffuse acoustic sounds, such as multi-channel crosstalk noise. As can be seen by comparing transfer functions T3 and T4, a directional microphone effectively suppresses diffuse acoustic sounds by approximately 5.5 dB compared to acoustic sounds with directional characteristics. This illustrates that a directional microphone is more sensitive to a speaker standing in front of the wearer of the speech intelligibility enhancement system 101 than it is to multi-channel crosstalk noise present in the room.
[0110] 8 and 9 show corresponding phase and delay plots of transfer functions such as those shown in FIG. 7. In FIG. 8, four phase curves P1 (squares), P2 (crosses), P3 (triangles), and P4 (circles) are shown, corresponding to the four transfer functions T1, T2, T3, and T4, respectively. The graph in FIG. 8 shows phase in degrees as a function of frequency in Hz. In FIG. 9, four group delay curves D1 (squares), D2 (crosses), D3 (triangles), and D4 (circles) are shown, corresponding to the four transfer functions T1, T2, T3, and T4, respectively. The graph in FIG. 9 shows delay in microseconds as a function of frequency in Hz. To illustrate the ability of the present invention to accommodate processing latency, a fixed delay of 100 microseconds has been added to the transfer functions of the electro-acoustic path.
[0111] Having identified the types of acoustic signals present in a room during a cocktail party situation (see FIG. 6) and the transfer functions of at least one in-ear headphone device 102 of a speech intelligibility enhancement system (see FIG. 7), the effect of applying said transfer functions to these signals will be considered with reference to the following figures.
[0112] FIG. 10 illustrates the effect of applying a vent transfer function T1 to three input audio signals represented by signal curves S1, S2, and S3. Similar to the graph of FIG. 6, the graph of FIG. 10 shows amplitude spectral density (ASD) as a function of frequency. Signal curve S4 shows the result of applying the vent transfer function T1 to the acoustic sound signal represented by signal curve S1. In other words, signal curve S4 shows the contribution of the vent / acoustic path to the multi-layer crosstalk noise present in the ear canal of a wearer of the in-ear headphone device 102. Signal curve S5 shows the result of applying the vent transfer function T1 to the acoustic sound signal represented by signal curve S2. In other words, signal curve S5 shows the contribution of the vent to the sound of a particular person speaking in the ear canal of a wearer of the in-ear headphone device. Signal curve S6 shows the result of applying the transfer function T1 to the acoustic signal represented by signal curve S3. In other words, signal curve S6 shows the contribution of the vent to the consonants produced by a particular person speaking and present in the ear canal of a wearer of an in-ear headphone device. As can be seen in Figure 10, the effect of the vent transfer function T1 is that the consonant (in this case, the consonant "t") is suppressed when passing through the vent with respect to lower frequency content.
[0113] FIG. 11 illustrates the effect of applying transfer function T3 to three input audio signals represented by signal curves S1, S2, and S3. Similar to the graph of FIG. 6, the graph of FIG. 11 shows amplitude spectral density (ASD) as a function of frequency. Signal curve S7 shows the result of applying transfer function T3 to the acoustic sound signal represented by signal curve S1. In other words, signal curve S7 shows the contribution of a directional microphone to the multi-layer crosstalk noise present in the ear canal of a wearer of the in-ear headphone device 102. Signal curve S8 shows the result of applying transfer function T3 to the acoustic sound signal represented by signal curve S2. In other words, signal curve S8 shows the contribution of a directional microphone to the desired audio signal present in the ear canal of a wearer of the in-ear headphone device. Signal curve S9 shows the result of applying transfer function T3 to the acoustic sound signal represented by signal curve S3. In other words, signal curve S9 shows the contribution of a directional microphone to the consonant "t" present in the ear canal of a user wearing an in-ear headphone device.
[0114] FIG. 12 is also a graph showing amplitude spectral density (ASD) as a function of frequency, similar to the graph in FIG. 6. This graph shows three signal curves, S10, S11, and S12. Signal curve S10 corresponds to the long-term average spectrum of the multiplexed crosstalk noise, which is equivalent to signal curve S1 in FIG. 6. Signal curve S11 corresponds to signal curve S4 seen in FIG. 9; that is, signal curve S11 shows the effect of applying a vent transfer function to the multiplexed crosstalk noise. Comparing signal curves S10 and S11 clearly shows the effect of the vent on the acoustic path, particularly the inherent low-pass characteristic of the vent. Low frequencies, i.e., multiplexed crosstalk noise with most of its energy in the vent's passband, pass through the vent, while higher frequencies are clearly attenuated by the presence of the vent. It can also be seen that the presence of the vent did not significantly reduce the amplitude spectral density for frequencies below 600 Hz.
[0115] However, signal curve S12 illustrates the effect obtained when vent transfer function T1 and omnidirectional microphone transfer function T2 (see FIG. 7) are applied to the multiple crosstalk noise signal S10 and combined in the ear canal. Essentially, signal curve S12 represents the long-term average spectrum of the multiple crosstalk noise present in the ear canal of a user of in-ear headphone device 102. As can be seen from this figure, the multiple crosstalk noise is significantly reduced compared to the multiple crosstalk noise without the in-ear headphone device (see signal curve S10). As can be seen from these exemplary transfer functions, a reduction of approximately 6 dB is achieved at 300 Hz.
[0116] FIG. 13 is also a graph showing amplitude spectral density (ASD) as a function of frequency, similar to the graph in FIG. 6 . Specifically, the graph shows three signal curves S13, S14, and S15. Signal curve S13 corresponds to the long-term average spectrum of a single speaker located approximately one meter away from the wearer of the speech intelligibility enhancement system 101; i.e., signal curve S13 corresponds to signal curve S2 as seen in FIG. 6 . Signal curve S14 shows the effect of applying vent transfer function T1 to the long-term average spectrum of the speaker's speech. Thus, signal curve S14 directly corresponds to signal curve S5 in FIG. 10 . Signal curve S15 represents the resulting audio signal present in the ear canal of a user wearing the in-ear headphone device 102, thereby representing contributions from the acoustic and electro-acoustic paths of the in-ear headphone device.
[0117] FIG. 14 also illustrates a graph of amplitude spectral density (ASD) as a function of frequency, similar to the graph in FIG. 6 . Specifically, the graph shows three signal curves S16, S17, and S18. Signal curve S16 corresponds to signal curve S3 in FIG. 6 and thus represents the short-term average spectrum of the consonant “t” produced by a speaker standing approximately one meter from the wearer of the in-ear headphone device 102. Signal curve S17 illustrates the effect of applying vent transfer function T1 to the short-term average spectrum of the consonant “t”; as can be seen in FIG. 14 , the vent significantly attenuates the signal. This is not surprising when considering the vent transfer function T1, which has a low-pass characteristic. Signal curve S18 illustrates the resulting consonant signal present in the ear canal of a user wearing the in-ear headphone device, thereby representing contributions from the acoustic and electro-acoustic paths of the in-ear headphone device. In this example, consonant amplification is achieved (as is evident when comparing signal curve S18 with signal curve S16). Such amplification is advantageous in that it further improves speech intelligibility, as is evident from the figures below.
[0118] FIG. 15 is also a graph showing amplitude spectral density (ASD) as a function of frequency, similar to the graph in FIG. 6 . Specifically, the graph shows four signal curves S19, S20, S21, and S22. Signal curve S19 corresponds to the multiple crosstalk noise signal, also seen as signal curve S1 in FIG. 6 , and signal curve S20 corresponds to the consonant signal, also seen as signal curve S3 in FIG. 6 . When directly comparing signal curves S19 and S20, it can be seen that when a user is not wearing an in-ear headphone device of a speech intelligibility enhancement system, low-frequency multiple crosstalk noise is present at a high level compared to the consonant “t.” The multiple crosstalk noise has a masking effect on the consonant “t” in this example. Note that although this example only concerns the letter “t,” a similar (and even more pronounced) effect often exists for other consonants. This makes speech understanding particularly difficult, as the consonant produced by the speaker of interest is “drowned out” by the multiple crosstalk noise present by other people in the room. Signal curve S21 corresponds to signal curve S12 as seen in FIG. 12, and signal curve S22 corresponds to signal curve S18 in FIG. 14. While signal curves S19-S22 have been effectively discussed in the preceding figures, the beneficial effect can first be truly appreciated when they are directly compared as in FIG. 15. As can be seen, signal curve S21 is reduced relative to signal curve S19 in the low-frequency range of the spectrum, i.e., in the vowel-dominant frequency range. Effectively, this indicates that the electrical acoustic pathway is arranged (by a specific transfer function as) to compensate for contributions from acoustic pathways / vents in the vowel-dominant frequency range, so that these contributions have a lower masking effect on consonants in the consonant-dominant frequency range. Improved speech intelligibility is thereby achieved. Thus, FIG. 15 shows that the signal-to-masking ratio is improved by the electrical acoustic pathway compensating for contributions from acoustic pathways in the vowel-dominant frequency range.
[0119] FIG. 16 also shows a graph of amplitude spectral density (ASD) as a function of frequency, similar to the graph in FIG. 6. Specifically, the graph shows four signal curves S23, S24, S25, and S26. Signal curve S23 corresponds to signal curve S1 (see FIG. 6), signal curve S24 corresponds to signal curve S2 (see FIG. 6), signal curve S25 corresponds to signal curve S12 (see FIG. 12), and signal curve S26 corresponds to signal curve S15 (see FIG. 13). This graph also reveals a beneficial effect on the long-term average spectrum of speech. When the in-ear headphone device 102 of the speech intelligibility enhancement system 101 is not used, the multi-pass crosstalk noise spectrum exceeds the long-term average spectrum of speech at high frequencies, such as frequencies in the range of 1 kHz to 6 kHz. However, when inserted into the ear canal, the speech intelligibility enhancement system 101 improves the speech-to-masker energy ratio, especially in the high-frequency range. This has the advantage that audio cues from the speaker of interest are easily detectable by the user of the system, thereby positively impacting speech intelligibility.
[0120] FIG. 17 illustrates an in-ear headphone device 102 of a speech intelligibility enhancement system according to an embodiment of the present invention. The in-ear headphone device 102 is arranged to apply transfer functions T1-T4 as illustrated in FIG. 7. Thereby, all of the signal processing results as illustrated throughout FIGS. 6-16 can be achieved by using the in-ear headphone device 102. The in-ear headphone device of this embodiment comprises two microphones 103 arranged to record acoustic sounds present in the external environment. The microphones in this embodiment are two omnidirectional microphones that are combined using a signal processor 104 to achieve the desired directional characteristics; however, dedicated directional microphones may also be used according to another embodiment of the present invention.
[0121] It should be noted that the voice intelligibility enhancement system 101 as mentioned in any of the foregoing descriptions may include two in-ear headphone devices 102, i.e., one in-ear headphone device 102 for each ear of the wearer of the voice intelligibility enhancement system. [Explanation of symbols]
[0122] 101 Voice Clarity Enhancement System 102 In-ear headphone device 103 Microphone 104 Signal Processor 105 loudspeaker 106 Loudspeaker duct 107 Vent 108 External acoustic environment 109 External auditory canal 110 Auricle (external ear) 111 Flexible Eartips 201 Eardrum (tympanic membrane) 202 Vent element 203 Damping Elements 204 Feedback Microphone 501 Acoustic Path 502 Electroacoustic Path 503 Sound pressure level difference 504 Center frequency of the vowel dominant frequency range 505 Center frequency of the consonant dominant frequency range 506 Resulting sound pressure level contributed by VDF 507 Resulting sound pressure level contributed by CDF VDF Vowel Dominant Frequency Range CDF Consonant Dominant Frequency Range S1~S26 signal curve T1~T4 transfer functions P1~P4 phase curve D1~D4 delay curve
Claims
1. 1. A speech intelligibility enhancement system for difficult acoustic conditions, the speech intelligibility enhancement system comprising at least one in-ear headphone device for insertion into a person's ear canal, the at least one in-ear headphone device being arranged with a portion facing the ear canal and a portion facing an environment, the at least one in-ear headphone device comprising: an acoustic pathway comprising a vent, the acoustic pathway connecting a portion facing the environment with a portion facing the ear canal; an electroacoustic path comprising a microphone in a portion facing the environment, a filter, and a loudspeaker in a portion facing the ear canal; the acoustic path is arranged to transmit acoustic sounds within a vowel dominant frequency range, and the electrical acoustic path is arranged to acoustically reproduce sound signals within a consonant dominant frequency range and within the vowel dominant frequency range; 1. A speech intelligibility enhancement system, wherein the electroacoustic path is arranged such that a signal-to-masking ratio is improved by the electroacoustic path compensating for contributions from the acoustic path in the vowel dominant frequency range.
2. 2. The speech intelligibility enhancement system of claim 1, wherein improving the signal-to-masking ratio comprises increasing the resulting sound pressure level present in the consonant-dominant frequency range relative to the resulting sound pressure level present in the vowel-dominant frequency range.
3. 3. The speech intelligibility enhancement system of claim 1, wherein the improving the signal-to-masking ratio comprises reducing a difference between a resulting sound pressure level in the ear canal contributed by the vowel dominant frequency range and a resulting sound pressure level in the ear canal contributed by the consonant dominant frequency range by using the electrical acoustic path to compensate for contributions from the acoustic path in the vowel dominant frequency range.
4. A speech intelligibility enhancement system according to any one of claims 1 to 3, wherein the consonant dominant frequency range includes frequencies above the vowel dominant frequency range.
5. 5. The speech intelligibility enhancement system of claim 1, wherein the acoustic path is arranged with an acoustic transfer function having a low-pass characteristic having a pass band and a cut-off frequency, the vowel dominant frequency range including frequencies below the cut-off frequency, and the consonant dominant frequency range including frequencies above the cut-off frequency.
6. The speech intelligibility enhancement system according to any one of claims 1 to 5, wherein the cut-off frequency is in the range of 250 Hz to 4 kHz.
7. The speech intelligibility enhancement system of any one of claims 1 to 6, wherein the vowel dominant frequency range comprises frequencies in the range of 50 Hz to 1 kHz.
8. A speech intelligibility enhancement system according to any one of claims 1 to 7, wherein the consonant dominant frequency range comprises frequencies in the range of 2 kHz to 4 kHz.
9. A speech intelligibility enhancement system according to any one of claims 1 to 8, wherein the difference is less than 15 dB, such as less than 10 dB, such as less than 8 dB, such as less than 6 dB, for example less than 5 dB.
10. A speech intelligibility system according to any one of the preceding claims, wherein the electro-acoustic path is arranged to compensate for contributions from the acoustic path within a signal processing frequency range of 300Hz to 1kHz.
11. A speech intelligibility enhancement system according to any one of claims 1 to 10, wherein the compensation of contributions from the acoustic paths is signal dependent.
12. A speech intelligibility enhancement system according to any one of claims 1 to 11, wherein the compensation for contributions from the acoustic paths is level dependent.
13. 13. The speech intelligibility enhancement system of claim 1, wherein the electrical acoustic path is arranged to compensate for contributions from the acoustic path by reproducing sound signals in at least a part of the vowel dominant frequency range.
14. 14. The speech intelligibility enhancement system of claim 1, wherein the electrical acoustic path is arranged to reproduce sound signals in the at least part of the vowel dominant frequency range having a polarity opposite to that of the acoustic sounds transmitted by the acoustic path. The effect of the compensation is that the perceived loudness of acoustic sounds in the vowel dominant frequency range is reduced compared to a situation where at least one in-ear headphone device is not inserted into the ear canal of the user / wearer.
15. 15. The speech intelligibility enhancement system of claim 1, wherein the electrical acoustic path is arranged to reproduce sound signals in the at least part of the vowel dominant frequency range by applying a phase shift to the sound signals.
16. 16. The speech intelligibility enhancement system of claim 15, wherein the phase shift is greater than 90 degrees and less than 270 degrees.
17. A speech intelligibility enhancement system according to any preceding claim, wherein the vent is an attenuated vent.
18. 18. The speech intelligibility enhancement system of any one of claims 1 to 17, wherein the loudspeaker and the vent are acoustically separated within the at least one in-ear headphone device.
19. 19. A speech intelligibility enhancement system according to any one of the preceding claims, wherein the vent is arranged with a cross-sectional area equivalent to a cylinder having a diameter in the range 1.5mm to 3.5mm, such as 2.0mm to 3.0mm, for example 2.3mm or 2.5mm.
20. 20. The speech intelligibility enhancement system of any one of claims 1 to 19, wherein the vent is arranged with a length equivalent to a cylinder having a length in the range of 2.5mm to 10mm, such as 3.5mm to 9mm, such as 4.5mm to 8mm, for example 5mm or 7mm.
21. The speech intelligibility enhancement system according to any one of claims 1 to 20, wherein the filter is arranged in a signal processor, such as a digital signal processor, of the at least one in-ear headphone device.
22. 22. The speech intelligibility enhancement system of any one of claims 1 to 21, wherein the at least one in-ear headphone device is battery powered, such as powered by a rechargeable battery.
23. 23. The speech intelligibility enhancement system of any one of claims 1 to 22, wherein the at least one in-ear headphone device comprises two in-ear headphone devices, one for each ear canal of the person, the two in-ear headphone devices being arranged to coordinate settings between them.
24. The speech intelligibility enhancement system according to any one of claims 1 to 23, wherein the at least one in-ear headphone device comprises a feedback microphone in a part facing the ear canal.
25. 25. The speech intelligibility enhancement system of claim 24, wherein the electrical acoustic path is arranged to compensate for contributions from the acoustic path based on input provided by the feedback microphone.
26. A speech intelligibility enhancement system according to any one of claims 1 to 25, wherein the microphones of the electro-acoustic path are directional microphones.
27. The speech intelligibility enhancement system of any one of claims 1 to 26, wherein the electro-acoustic path of the at least one in-ear headphone device comprises a plurality of microphones.
28. A speech intelligibility enhancement system according to any preceding claim, wherein the electro-acoustic path is arranged to amplify sound with a nominal gain in a passband of the electro-acoustic path.
29. 1. A method for enhancing speech intelligibility in difficult acoustic conditions, the method comprising: inserting at least one in-ear headphone device into the person's ear canal, the at least one in-ear headphone device being arranged with an ear canal-facing portion and an environment-facing portion, the at least one in-ear headphone device comprising an acoustic path comprising a vent coupling the environment-facing portion with the ear canal-facing portion, and an electro-acoustic path comprising a microphone in the environment-facing portion, a filter, and a loudspeaker in the ear canal-facing portion; transmitting acoustic sounds within a vowel dominant frequency range through the acoustic pathway from a portion facing the environment to a portion facing the ear canal; acoustically reproducing, by said electroacoustic path, sound signals within a consonant dominant frequency range and within said vowel dominant frequency range; and compensating by said electro-acoustic path for contributions from said acoustic path in said vowel dominant frequency range so as to improve the signal to masking ratio.
30. The method of claim 29, wherein the method is performed by a speech intelligibility enhancement device according to any one of claims 1 to 28.