Speech enhancement with active masking control
By combining acoustic paths and electroacoustic paths in in-ear headphone devices, the problem of low voice intelligibility in noisy environments is solved, and the effect of improving signal masking ratio and improving voice intelligibility is achieved.
Patent Information
- Application Number
- CN202380076512.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-31
- Filing Date
- 2023-10-30
- Publication Date
- 2025-06-13
AI Technical Summary
In noisy acoustic environments, it is difficult for people with normal hearing or close to normal hearing to understand speech, and the prior art noise suppression algorithms and hearing protectors cannot effectively improve speech intelligibility.
A voice intelligibility enhancement system is designed, including in-ear headphone devices, which include acoustic paths and electroacoustic paths. The acoustic path passes ambient acoustic sound to the ear canal through the vent, the electroacoustic path reproduces the sound signal through the microphone, filter and speaker, and compensates for the acoustic path contribution in the vowel dominant frequency range through signal processing, improving the signal masking ratio.
By improving the signal masking ratio, the difference between the sound pressure level in the vowel dominant frequency range and the sound pressure level in the consonant dominant frequency range is reduced, and the intelligibility of speech in noisy environments is improved and listening comfort is enhanced.
Smart Images

Figure CN120153669A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a speech intelligibility enhancement system for difficult acoustic conditions and a method for enhancing speech intelligibility under difficult acoustic conditions. Background Art
[0002] Speech communication difficulties in noisy environments are a common experience. Cocktail parties, cafes, and similar situations in particular pose challenges because the signals (the speech of the conversation partner) are very similar and usually not as loud as the noise (the speech of other people). People with normal hearing need to expend a great deal of effort to distinguish the discourse, and even people with very mild hearing loss need to expend more effort.
[0003] Many noise suppression algorithms, including adaptive microphone patterns, exhibit significant gains in signal-to-noise ratio (SNR). However, they generally do not provide better speech recognition scores in practical tests, for example due to processing artifacts and unnatural sounds.
[0004] Traditional passive hearing protectors usually attenuate too much, especially at higher frequencies, making speech recognition even worse. In addition, traditional hearing protectors cause occlusion (i.e., without compensation due to ear canal occlusion, the user perceives a "hollow" or "booming" sound of their own voice).
[0005] So-called musician earplugs, which are designed to attenuate a wide audio frequency band relatively evenly without distorting the perception of music, usually also attenuate too much to be used for listening in noisy environments. They also often do not address the occlusion effect.
[0006] On the other hand, hearing aids are designed to improve audibility by using general measures of sound amplification. As described above, this generally does not help people with normal or near-normal hearing who have difficulty understanding speech in noisy environments. To allow the user to participate in conversations, most hearing aids contain vents to allow the bone / tissue-conducted sound of the user's own voice to escape from the ear canal, but this has the inherent problem that when the vent is large enough to provide an acceptable perception of the user's own voice, a lot of low-frequency energy from the surrounding environment enters the ear, is amplified by Helmholtz resonance, and masks important higher-frequency speech cues. To counteract this masking effect, the high-frequency gain must be increased. This in turn means that the overall level at the eardrum increases to a level higher than that which would result in an open ear. However, when the level increases above a certain level (for a normal hearing subject, corresponding to about 65 dBA outside the ear, this certain level is much lower than the level present at a typical party), frequency discrimination and speech understanding decline.
[0007] An ear device that addresses one or more of the above challenges of improving listening comfort and / or speech recognition in noisy environments for people with normal or near-normal hearing would be highly advantageous and useful. Summary of the Invention
[0008] The inventors have recognized the above problems and challenges, particularly those related to listening comfort and intelligibility of conversations in noisy environments, and have subsequently made the invention described below.
[0009] One aspect of the present invention relates to a speech intelligibility enhancement system for difficult acoustic conditions, the speech intelligibility enhancement system comprising at least one in-ear headphone device for insertion into a person's ear canal, the at least one in-ear headphone device being arranged with an ear canal-facing portion and an environment-facing portion, and the at least one in-ear headphone device comprising:
[0010] an acoustic path including a vent, the acoustic path coupling the environment-facing portion to the ear canal-facing portion; and
[0011] an electroacoustic path including a microphone, a filter at the environment-facing portion, and a speaker at the ear canal-facing portion;
[0012] wherein the acoustic path is arranged to transmit acoustic sounds in a vowel-dominated frequency range, and wherein the electroacoustic path is arranged to acoustically reproduce sound signals in a consonant-dominated frequency range and in the vowel-dominated frequency range; and
[0013] wherein the electroacoustic path is arranged to improve the signal-to-mask ratio by compensating for the contribution from the acoustic path in the vowel-dominated frequency range through the electroacoustic path.
[0014] Thus, an advantageous system for enhancing speech intelligibility in difficult acoustic environments is provided. The advantages of the system will become clear below.
[0015] In this context, speech is understood as sound communication using a language, such as a non-tonal language. Each language uses a phonetic combination of vowels and consonants that form its utterances. Vowels tend to be lower in frequency and louder than consonants, and thus carry the major part of the sound energy caused by speech. However, in fact, it is the lower-energy and higher-frequency consonants that carry most of the meaning of the utterance. Therefore, the intelligibility of speech highly depends on the frequency range of the speech caused by consonants. Compared with vowels, consonants are more sensitive to the upward spread of masking, and thus the energy from vowels can exert a masking effect on consonants. Such a masking effect is likely to occur because it may happen when listening to a person speaking in a loud acoustic environment, such as in a café with a high level of background noise. To overcome the high level of background noise, people tend to speak louder and need more effort to hear, a phenomenon commonly known as the Lombard effect. Speaking louder has a profound impact on the intelligibility for others because the added sound energy is concentrated around the vowels, i.e., in the vowel-dominated frequency range, while only very little energy can be added to the consonants, i.e., in the consonant-dominated frequency range. Therefore, in a social environment where everyone speaks louder, it means that a large amount of energy is added in the vowel-dominated frequency range. This essentially means that the consonants will be masked by the vowels, which in turn makes it more difficult to understand what is being said. However, the importance of vowels in speech should not be underestimated, and they still play an important role in speech intelligibility.
[0016] In this context, the signal-to-masking ratio is understood as a measure that compares the level of a desired signal with the level of a masking signal. In this context, the desired signal is a signal that is substantially present in the consonant-dominated frequency range, and the masking signal is a signal that is substantially present in the vowel-dominated frequency range. In other words, the signal-to-masking ratio can also be referred to as the consonant-to-vowel ratio. The masking signal may not necessarily represent the unwanted sound typical for a noise signal when discussing the signal-to-noise ratio. Instead, the masking signal can actually also include speech cues that contribute to speech intelligibility. For example, the masking signal can include the sound contribution of the speaker of interest (the person speaking to the wearer of the speech intelligibility enhancement system), as well as the sound contributions made by multiple other people present in the same acoustic environment as the speaker of interest and the wearer of the speech intelligibility enhancement system (throughout the following disclosure, this sound contribution may be referred to as babble noise). The point is that the masking signal exerts a masking effect on the consonants in the consonant-dominated frequency range, and thus, by improving the signal-to-masking ratio, the masking effect can be reduced, and speech intelligibility can be improved. Therefore, it should be noted that the speech intelligibility enhancement system is thus effectively arranged to perform active masking control.
[0017] It should also be noted that the foregoing discussion regarding improving the signal-to-mask ratio should be interpreted as an improvement based on the situation where the user is not wearing the speech intelligibility enhancement system, that is, based on the situation where at least one in-ear headphone device (such as two in-ear headphone devices) is not inserted into the user's ear canal.
[0018] According to one embodiment, improving the signal-to-mask ratio includes increasing the resulting sound pressure level present in the consonant-dominated frequency range relative to the resulting sound pressure level present in the vowel-dominated frequency range.
[0019] Improving the signal-to-mask ratio or the consonant-to-vowel ratio may include increasing the resulting sound pressure level in the consonant-dominated frequency range relative to the resulting sound pressure level present in the vowel-dominated frequency range. This may include amplifying the acoustic sound present in the consonant-dominated frequency range, that is, the electroacoustic path is arranged to perform sound amplification in the consonant-dominated frequency range. However, this should not be interpreted as achieving an improved signal-to-mask ratio only by adjusting the gain of the electroacoustic path in the consonant-dominated frequency range, because the electroacoustic path is still arranged to compensate for the contribution of the acoustic path in the vowel-dominated frequency range. Increasing the resulting sound pressure level present in the consonant-dominated frequency range relative to the resulting sound pressure level present in the vowel-dominated frequency range is advantageous because the effect of the masking signal on the signal of most interest for speech intelligibility (i.e., the signal in the consonant-dominated frequency range) is reduced, thereby improving speech intelligibility.
[0020] According to one embodiment, improving the signal-to-mask ratio includes: reducing the difference between the resulting sound pressure level in the ear canal contributed by the vowel-dominated frequency range and the resulting sound pressure level in the ear canal contributed by the consonant-dominated frequency range by using the electroacoustic path to compensate for the contribution from the acoustic path in the vowel-dominated frequency range.
[0021] The speech intelligibility enhancement system according to this embodiment is advantageous because it reduces the difference between the sound pressure level (SPL) contributed by the vowel-dominated frequency range and the SPL contributed by the consonant-dominated frequency range. The sound pressure level is the most commonly used measure of the intensity of a sound wave and is typically measured in decibels (dB). The reduction of the sound pressure level difference is performed by using an electroacoustic path to compensate for the contribution from the acoustic path. The purpose of the compensation is not to completely cancel out the contribution from the acoustic path, because the vowel content of the speech contributed by the acoustic path is still important in the reproduction of speech in the ear canal of a person wearing an in-ear headphone device. Without the vowel content contributed by the acoustic path, the speech as experienced by the wearer of at least one in-ear headphone device would sound unnatural and lack important characteristics. However, the purpose of the compensation is to reduce the influence of high-energy vowels relative to the influence of lower-energy consonants. Thus, the consonant-dominated part of the speech can be enhanced relative to the vowel-dominated part of the speech, thereby improving the intelligibility of the speech in many acoustic environments.
[0022] In other words, the effect of the compensation is that the total transfer function from the external acoustic environment to the ear canal, which is produced by the contributions of both the acoustic path and the electroacoustic path, exhibits a smaller difference between the sound pressure level in the vowel-dominated frequency range and the sound pressure level in the consonant-dominated frequency range compared to the difference between the two without the application of the compensation.
[0023] According to an embodiment of the present invention, the reduction of the difference is obtained by using a filter implemented in a signal processor to compensate for the contribution from the acoustic path. The signal processor can apply filtering to the signal recorded by the microphone, thereby providing a filtered signal for reproduction using the speaker. The effect of the acoustic reproduction of the filtered signal is that the effect of the acoustic sound contributed by the acoustic path in a sub-range or the full range of the vowel-dominated frequency range is attenuated.
[0024] In this context, the vowel-dominated frequency range is understood to be a frequency range that is substantially dominated by the presence of frequency components that form the vowel part. Furthermore, in this context, the consonant-dominated frequency range is understood to be a frequency range that is dominated by the presence of frequencies that form the consonant part. Those skilled in the art will readily understand that it is not possible to draw a clear boundary between the frequency components of vowels and the frequency components of consonants, because any tone generated by a human can include multiple harmonics, including the fundamental harmonic (or fundamental wave) and the second harmonic, third harmonic, fourth harmonic, etc. (or overtones), and the overtones of the frequency components of vowels can exist in higher frequency ranges, such as in the consonant-dominated frequency range. However, those skilled in the field of speech will understand that the frequency components of consonants generally exist at higher frequencies (such as from 2 kHz to 4 kHz) than the frequency components that make up vowels, and vowels generally exist at lower frequencies (such as in the range from 50 Hz to 1 kHz).
[0025] At least one in-ear headphone device (such as two in-ear headphone devices) of a voice intelligibility enhancement system (or the "system" hereinafter) is arranged to be inserted into a person's ear canal. When inserted into the ear canal, the in-ear headphone device has a part facing the ear canal - the "ear canal-facing part" - and a part facing another direction of the person's surrounding environment - the "environment-facing part". These two parts of the in-ear headphone device are coupled by the presence of an acoustic path. The acoustic path is understood as a path along which acoustic sound can propagate. The acoustic path includes a vent, which is a channel or duct having a specific geometry that can be determined by acoustic problems. The vent effectively couples the environment-facing part to the ear canal-facing part, ensuring that acoustic sound present in the environment can propagate into the person's ear canal. The acoustic path is arranged such that acoustic sound in a specific frequency range can propagate through the acoustic path, while acoustic sound of other frequencies may be impeded. These acoustic characteristics can be attributed to the geometry of the vent (shape, cross-sectional area, length). Generally, in systems of the prior art, such as in-ear headphone devices for listening to music, such vents are used to reduce the impact of the occlusion effect on the listening experience. However, as will be clear hereinafter, the presence of the vent is for another purpose, namely the acoustic reproduction of sound in a person's ear canal. Nevertheless, the beneficial effect of the presence of the vent is that the acoustic path can reduce the impact of the occlusion effect on the wearer's experience of their own voice.
[0026] In addition to having an acoustic path including a vent, the in-ear headphone device further includes an electroacoustic path, which includes a microphone, a filter at the environment-facing part, and a speaker at the ear canal-facing part. By means of such an electroacoustic path, sound from the external acoustic environment can be electronically processed, for example, digitally processed, and reproduced in the ear canal.
[0027] Furthermore, the advantage of the voice intelligibility enhancement system is that it can effectively provide multi-band (such as dual-band) dynamic range compression, thereby achieving different compressions in the vowel-dominated range and the consonant-dominated range.
[0028] In addition, the voice intelligibility enhancement system is advantageous because it can at least to some extent provide a natural reproduction of the acoustic sound in the external environment. This effect is provided at least by the acoustic path, which enables the natural reproduction in the ear canal of the sound present in the external acoustic environment.
[0029] An in-ear headphone device can be understood as a headphone device arranged to be worn by a user by fitting the device in the user's outer ear (such as in the concha, near the ear canal). The in-ear headphone device can also at least partially extend into the user's ear canal. The in-ear headphone device can generally be shaped to at least partially fit within the outer ear and / or ear canal, thereby ensuring that the device fits into the user's ear. The in-ear headphone device can also be understood as an in-ear headphone, earplug, canal headphone, earbud headphone, or audible wearing device.
[0030] According to one embodiment, the consonant-dominated frequency range includes frequencies higher than the vowel-dominated frequency range.
[0031] The consonant-dominated frequency range can include frequencies higher than the vowel-dominated frequency range. Regardless of whether there is an overlap between the consonant-dominated frequency range and the vowel-dominated frequency range, the consonant-dominated frequency range can still include frequencies that do not exist in the vowel-dominated frequency range, and these frequencies are higher than the vowel-dominated frequency range. Thus, it is clear that the consonant-dominated frequency range is a frequency range associated with frequencies higher than the vowel-dominated frequency range.
[0032] According to one embodiment, the acoustic path is arranged with an acoustic transfer function that has a low-pass characteristic, the low-pass characteristic having a passband and a cut-off frequency, and wherein the vowel-dominated frequency range includes frequencies lower than the cut-off frequency, and wherein the consonant-dominated frequency range includes frequencies higher than the cut-off frequency.
[0033] The acoustic path can be arranged in such a way that it has an acoustic transfer function that has a low-pass characteristic, the low-pass characteristic having a passband and a cut-off frequency, where the vowel-dominated frequency range includes frequencies lower than the cut-off frequency, and where the consonant-dominated frequency range includes frequencies higher than the cut-off frequency. By implementing such a low-pass characteristic in its transfer function, the acoustic path is arranged to allow acoustic sounds with frequencies lower than the cut-off frequency (in the passband) to pass through and reach the ear canal of a user wearing at least one in-ear headphone device. However, acoustic sounds with frequencies higher than the cut-off frequency are severely restricted when passing through the acoustic path. Those skilled in the art of acoustics will readily understand that acoustic sounds with frequencies higher than the cut-off frequency can pass along the acoustic path, however, they are severely impeded. Generally, the cut-off frequency is the frequency at which 3 dB of attenuation occurs. The vowel-dominated frequency range includes frequencies lower than the cut-off frequency, and the consonant-dominated frequency range includes frequencies higher than the cut-off frequency, however, this does not exclude the possibility that the vowel-dominated frequency range may also include frequencies higher than the cut-off frequency and the consonant-dominated frequency range may include frequencies lower than the cut-off frequency.
[0034] According to one embodiment, the cut-off frequency is in the range of 250 Hz to 4 kHz.
[0035] The cut-off frequency can be in the range of 250 Hz (Hertz) to 4 kHz (kilohertz), such as 500 Hz to 2 kHz, such as 650 Hz to 1600 Hz, such as 700 Hz to 1200 Hz, such as 800 Hz, 900 Hz or 1 kHz.
[0036] According to one embodiment, the vowel-dominant frequency range includes frequencies in the range of 50 Hz to 1 kHz.
[0037] The vowel-dominant frequency range can include frequencies in the range from 50 Hz to 1 kHz, such as frequencies in the range from 400 Hz to 800 Hz, such as 600 Hz.
[0038] According to one embodiment, the consonant-dominant frequency range includes frequencies in the range of 2 kHz to 4 kHz.
[0039] According to an embodiment, the difference is less than 15 dB, such as less than 10 dB, such as less than 8 dB, such as less than 6 dB, for example less than 5 dB.
[0040] After compensation, the difference between the sound pressure level in the ear canal contributed by the vowel-dominant frequency range and the resulting sound pressure level in the ear canal contributed by the consonant-dominant frequency range can be less than 15 dB (decibels), such as less than 10 dB, such as less than 8 dB, such as less than 6 dB, for example less than 5 dB.
[0041] Reducing the difference between the sound pressure level contributed by the vowel-dominant frequency range and the sound pressure level contributed by the consonant-dominant frequency range such that the difference is less than 15 dB is considered to only reduce the impact provided by the acoustic path on the listening experience, rather than completely eliminate the impact provided by the acoustic path. In this context, the acoustic sound propagated through the acoustic path is not considered an unwanted sound, and on the contrary, the acoustic sound propagated from the ambient-facing part of the in-ear headphone device to the ear canal-facing part of the in-ear headphone device through the acoustic path can contribute to speech intelligibility.
[0042] According to one embodiment, the electroacoustic path is arranged to compensate for the contribution from the acoustic path in the signal processing frequency range of 300 Hz to 1 kHz.
[0043] The electroacoustic path can be arranged to compensate for the contribution from the acoustic path in the signal processing frequency range by using a signal processor. The signal processing frequency range can be a frequency range of 300 Hz to 1 kHz, such as a frequency range of 400 Hz to 800 Hz, such as 600 Hz.
[0044] According to one embodiment, the compensation of the contribution from the acoustic path is signal-dependent.
[0045] Signal-dependent is at least understood as that the signal processing (i.e., compensation) depends on the acoustic signals present in the external acoustic environment. Such signal-dependent compensation is advantageous because the speech intelligibility enhancement system can better adapt to the external acoustic environment, thereby providing an improved listening experience.
[0046] According to one embodiment, the compensation of the contribution from the acoustic path is level-dependent.
[0047] It should be understood that the signal processing (i.e., compensation) depends on the sound pressure level measured, for example, in the vowel-dominated frequency range (e.g., the center frequency of the vowel-dominated frequency range). The advantage of such a dependence is that the speech intelligibility enhancement system can better adapt to the sound pressure level present in the external acoustic environment. For example, if the sound pressure level in the external acoustic environment is low, the compensation requirements for achieving a high level of speech intelligibility may be less compared to the case of a high sound pressure level. In such a low sound pressure level, the compensation can be kept at a low level, thus affecting the acoustic sound contributed by the acoustic path less severely, e.g., with less distortion. The following advantage is thus achieved: the reproduction quality of the acoustic sound is always as high as possible in the ear canal of a user wearing at least one in-ear headphone device.
[0048] It should also be noted that the device implementing any of the above clauses can be arranged to intermittently adjust the compensation according to changes in the external acoustic environment, including adjusting the gain of the transfer function and even turning the compensation on and off. The device can additionally be arranged to perform other types of sound processing according to other acoustic conditions. Such other types of processing can include low-frequency amplification, which may be advantageous in quiet conversations.
[0049] According to one embodiment, the electroacoustic path is arranged to compensate for the contribution from the acoustic path by reproducing a sound signal in at least a part of the vowel-dominated frequency range.
[0050] The electroacoustic path can be arranged to reproduce the sound signal of the external acoustic environment in at least a part of the vowel-dominated frequency range. This can, for example, include using the speaker of the in-ear headphone device to reproduce the sound signal in a way that compensates for the contribution from the acoustic path. The compensation can include reproducing a sound signal with a polarity opposite to or a different phase from the audio signal contributed by the acoustic path in at least the vowel-dominated frequency range.
[0051] According to one embodiment, the electroacoustic path is arranged to reproduce a sound signal in at least a portion of the vowel-dominated frequency range with a polarity opposite to the polarity of the acoustic sound transmitted by the acoustic path. The effect of the compensation is that the perceived loudness of the acoustic sound in the vowel-dominated frequency range is reduced compared to the case where at least one in-ear headphone device is not inserted into the ear canal of the user / wearer.
[0052] According to one embodiment, the electroacoustic path is arranged to reproduce a sound signal in at least a portion of the vowel-dominated frequency range by applying a phase shift to the sound signal.
[0053] Since the electroacoustic path includes a microphone or multiple microphones according to other embodiments and a speaker, the electroacoustic path can perform signal processing on the recorded signal. The signal processing can include applying a phase shift to such a signal. Thus, the electroacoustic path can be arranged to reproduce, by applying a phase shift, a sound signal originating from the external acoustic environment in the ear canal of the user in at least a portion of the vowel-dominated frequency range. The application of the phase shift can cause the reproduced signal to have a canceling effect on the audio signal transmitted from the external acoustic environment to the ear canal via the acoustic path and its vent.
[0054] According to an embodiment, the phase shift is higher than 90 degrees and lower than 270 degrees.
[0055] The applied phase shift can be higher than 90 degrees and lower than 270 degrees. The phase shift can be applied to any frequency in the vowel-dominated frequency range, such as the center frequency of the vowel-dominated frequency range, for example, at a frequency of 600 Hz.
[0056] According to an embodiment of the present invention, the microphone and the speaker are wired oppositely with respect to the positive terminal and the negative terminal.
[0057] According to one embodiment, the vent is a damped vent.
[0058] The vent can be a damped vent including one or more vent elements and one or more damping elements. The damped vent can be, for example, a vent having a damping cloth located at one or both ends of the vent, or a vent configured with an integrated damping effect.
[0059] When the user wears the in-ear headphone device, adding an undamped vent to the in-ear headphone device can suppress the occlusion effect but cause Helmholtz resonance. By further adding a damping element to the vent, thereby providing a damped vent, the Helmholtz resonance and its possible associated distortion can be removed.
[0060] According to one embodiment, the speaker and the vent are acoustically separated inside the at least one in-ear headphone device.
[0061] According to one embodiment, the vent is arranged to have the same cross-sectional area as a cylinder with a diameter in the range of 1.5 mm to 3.5 mm, such as 2.0 mm to 3.0 mm, for example 2.3 mm or 2.5 mm.
[0062] The preferred cross-sectional area of the vent can be, for example, in the range of 1.8 mm 2 (square millimeters) to 9.6 mm 2 , such as 3.1 mm 2 to 7.1 mm 2 , for example 4.2 mm 2 or 4.9 mm 2 . The vent can have various cross-sectional shapes, such as circular, rectangular, and semi-circular, and can have a varying cross-sectional area along its length, or be composed of two or more vents or branched vents, but is preferably designed to have dimensions equal to those of the cylindrical vent described above.
[0063] According to one embodiment, the vent is arranged to have a length equal to that of a cylinder with a length in the range of 2.5 mm to 10 mm, such as 3.5 mm to 9 mm, for example 4.5 mm to 8 mm, for example 5 mm or 7 mm.
[0064] The vent can have various shapes along its length and can be straight, curved, or bent, and can be composed of two or more vents or branched vents, but is preferably designed to have dimensions equal to those of the cylindrical vent described above.
[0065] According to one embodiment, the filter is arranged in a signal processor (such as a digital signal processor) of the at least one in-ear headphone device.
[0066] According to one embodiment, the at least one in-ear headphone device is battery-powered, for example, powered by a rechargeable battery.
[0067] According to one embodiment, the at least one in-ear headphone device includes two in-ear headphone devices, each ear canal of the person uses one in-ear headphone device, and the two in-ear headphone devices are arranged to coordinate their settings with each other.
[0068] A voice intelligibility enhancement system may include two in-ear headphone devices, one for each ear canal of a person, and the two devices are arranged to coordinate their settings. Thus, a voice intelligibility enhancement system is achieved that has the same advantages as described above and is suitable for use with both ears of a user simultaneously. It should be noted that any effects and advantages described with respect to at least one in-ear headphone device equally apply to the two in-ear headphone devices of this embodiment.
[0069] According to one embodiment, the at least one in-ear headphone device includes a feedback microphone at the ear canal-facing portion.
[0070] At least one in-ear headphone device of a voice intelligibility enhancement system may include a feedback microphone arranged at the ear canal-facing portion of the at least one in-ear headphone device. The advantage of the feedback microphone is that it helps to improve the control of the sound processing performed by the electroacoustic path of the at least one in-ear headphone device. In particular, the feedback microphone can be used to adapt the feedforward processing of the electroacoustic path. Additionally, the advantage of the feedback is that it enables the voice intelligibility enhancement system to detect whether the user / wearer of the system is speaking and accordingly adjust the electroacoustic path to provide the user / wearer with the desired impression of their own voice.
[0071] According to one embodiment, the electroacoustic path is arranged to compensate for the contribution from the acoustic path based on the input provided by the feedback microphone.
[0072] The advantage of compensating for the contribution from the acoustic path based on the input provided by the feedback microphone is the improved control of the sound processing performed by the electroacoustic path of the at least one in-ear headphone device. Specifically, by making the compensation based on the input provided by the feedback microphone, it can be ensured that the acoustic sound present in the user's ear canal when the at least one in-ear headphone device is inserted therein actually reflects the desired listening experience.
[0073] According to one embodiment, the microphone of the electroacoustic path is a directional microphone.
[0074] In a preferred embodiment of the present invention, the microphone in the electroacoustic path of at least one in-ear headphone device is a directional microphone. A directional microphone is understood to be a microphone that is most sensitive in one or more directions. In other words, a directional microphone has a polar pattern different from omnidirectional. Those skilled in the art will readily understand that such a directional microphone can be implemented in various ways, including using multiple microphones arranged in a specific configuration, or by using a single microphone in combination with multiple microphone ports / tubes. When implemented in at least one in-ear headphone device, the advantage of a directional microphone is that omnidirectional sound contributions (such as multi-path coincidence noise) can be suppressed relative to sound contributions with more directional characteristics (such as the relevant speech of a speaker standing in front of a user / wearer of a voice intelligibility enhancement system). Therefore, voice intelligibility can be further improved.
[0075] According to an embodiment of the present invention, the directional microphone has a supercardioid characteristic.
[0076] According to one embodiment, the electroacoustic path of the at least one in-ear headphone device includes multiple microphones.
[0077] According to an embodiment of the present invention, the electroacoustic path of at least one in-ear headphone device may include multiple microphones, for example, two microphones. The multiple microphones may be arranged such that at least one in-ear headphone device includes a directional microphone and an omnidirectional microphone.
[0078] According to one embodiment, the electroacoustic path is arranged to amplify sound with a nominal gain in the passband of the electroacoustic path.
[0079] The electroacoustic path may be arranged to amplify sound with a nominal gain in the passband of the electroacoustic path, such as amplifying sound with a nominal gain throughout the passband of the electroacoustic path. This is advantageous in cases of low sound pressure levels where voice understanding may be difficult.
[0080] Another aspect of the present invention relates to a method for enhancing voice intelligibility under difficult acoustic conditions, the method comprising the following steps:
[0081] Inserting at least one in-ear headphone device into a person's ear canal, the at least one in-ear headphone device being arranged with an ear canal-facing portion and an environment-facing portion, the at least one in-ear headphone device including an acoustic path and an electroacoustic path, the acoustic path including a vent that couples the environment-facing portion to the ear canal-facing portion, and the electroacoustic path including a microphone, a filter at the environment-facing portion, and a speaker at the ear canal-facing portion;
[0082] Transmitting acoustic sound in the vowel-dominated frequency range from the environment-facing portion to the ear canal-facing portion through the acoustic path;
[0083] Acoustically reproduce sound signals in the consonant-dominated frequency range and the vowel-dominated frequency range through the electroacoustic path; and
[0084] Compensate for the contribution from the acoustic path in the vowel-dominated frequency range through the electroacoustic path, so that the signal-to-mask ratio is improved.
[0085] Thus, a method for enhancing speech intelligibility under difficult acoustic conditions is achieved. This method is advantageous for at least the same reasons as those given for the above-mentioned speech intelligibility enhancement system.
[0086] According to an embodiment, the method is performed by a speech intelligibility enhancement device according to any of the foregoing clauses. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Various embodiments of the present invention will be described below with reference to the accompanying drawings, in which,
[0088] Figure 1 An in-ear headphone device of a speech intelligibility enhancement system according to an embodiment of the present invention is shown,
[0089] Figures 2a to 2d Various in-ear headphone devices according to embodiments of the present invention are shown,
[0090] Figures 3a to 3h Various layouts of air vents of an acoustic path suitable for an in-ear headphone device according to an embodiment of the present invention are shown,
[0091] Figure 4 The characteristics of an acoustic path and an electroacoustic path according to an embodiment of the present invention are shown,
[0092] Figure 5 The concept of compensating for the contribution from the acoustic path in the vowel-dominated frequency range using the electroacoustic path is shown,
[0093] Figures 6 to 16 The spectrum of a sound signal present in a room, the transfer function of an in-ear headphone device according to an embodiment of the present invention, and the application of the transfer function of a sound signal useful for understanding the present invention are shown, and
[0094] Figure 17 A speech intelligibility enhancement system according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0095] Figure 1FIG. 0 shows a speech intelligibility enhancement system 101 according to an embodiment of the present invention. The speech intelligibility enhancement system 101 is shown to include an in-ear headphone device 102. However, according to another embodiment, the speech intelligibility enhancement system 101 may include two in-ear headphone devices 102; one in-ear headphone device is used for each ear of a person. Therefore, the following description related to the in-ear headphone device 102 also applies to a system including two in-ear headphone devices.
[0096] Figure 1 FIG. 4 shows the in-ear headphone device 102 when inserted into the ear canal 109 of a person / user wearing the in-ear headphone device. The in-ear headphone device 102 is preferably placed in the outer ear 110 of the user and is provided with a flexible earplug 111 for providing an acoustic seal in the ear canals 109 of different users.
[0097] The in-ear headphone device 102 includes a microphone 103 arranged to primarily record acoustic sounds from the external acoustic environment 108. In the drawings of this embodiment, the microphone 103 is shown arranged at the end of the in-ear headphone device 102 facing the external acoustic environment. However, in other embodiments of the present invention, the microphone 103 may also be arranged inside the in-ear headphone device 102 and acoustically coupled to the external acoustic environment 108 through a microphone duct (not shown in the figure). The in-ear headphone device also includes a signal processor 104 in the form of a digital signal processor, which is configured to receive the recorded audio signal from the microphone 103 and apply a filter (a digital filter in this embodiment) thereto to provide a filtered audio signal for acoustic reproduction using the speaker 105 of the in-ear headphone device. In the drawings of this embodiment, the speaker 105 is shown to be contained within the in-ear headphone device 102, and the acoustic sound emitted by the speaker 105 is transmitted to the ear canal 109 via the speaker duct 106. However, in other embodiments, the speaker duct 106 may be omitted, and the speaker 105 may be arranged closer to the end of the in-ear headphone device 102 facing the ear canal. Hereinafter, the set including the microphone 102, the signal processor 104, and the speaker 105 is referred to as an electroacoustic path.
[0098] In addition to the electroacoustic path, the in-ear headphone device 102 also includes an acoustic path, which includes a vent 107. The vent is a narrow duct through which sound can propagate. The purpose of the vent is to facilitate the transmission of low-frequency acoustic sound between the ear canal 109 and the external acoustic environment 108. In other words, the vent 107 enables the coupling of the ambient-facing part of the in-ear headphone device 102 with the ear canal-facing part of the in-ear headphone device 102. The boundary between the ear canal-facing part and the ambient-facing part of the in-ear headphone device 102 is at the circumference of the in-ear headphone device 102, where it typically contacts the ear canal 109, i.e., where it substantially plugs the ear canal.
[0099] Figures 2a to 2d Various in-ear headphone devices 102 according to embodiments of the present invention are shown.
[0100] Figure 2a The in-ear headphone device 102 that is also inserted into the user's ear canal 109 according to an embodiment is shown. Figure 1 As shown, acoustic sound present in the external acoustic environment 108 can propagate through the acoustic path of the in-ear headphone device 102, i.e., through the vent 107 and its vent element 202, and into the user's ear canal 109. In addition, acoustic sound present in the external acoustic environment 108 is picked up by the microphone 103, processed by the signal processor 104, acoustically reproduced by the speaker 105, and the reproduced sound is guided from the speaker 105 to the ear canal 109 via the speaker duct 106. Thus, it is clear that the total transfer function of the sound from the external acoustic environment 108 and into the ear canal 109 includes two contributions, namely the acoustic path and the electroacoustic path. Therefore, the sound picked up by the user's eardrum 201 is produced by these contributions. As shown in this figure, the vent includes a single vent element 202 in the form of a duct. However, as will be clear from the following description, other configurations of the vent are possible according to other embodiments.
[0101] Figure 2b A variant of the in-ear headphone device 102 as shown according to another embodiment is shown. Figure 2a In this embodiment, the vent 107 is a damped vent, which additionally includes a damping element 203. The damping element according to the present embodiment is a damping cloth located at one end of the damped vent 107. In another embodiment, the damping characteristics of the damped vent 107 are provided by damping cloths at both ends of the damped vent 107, and in other embodiments, the damping characteristics of the damped vent 107 are provided by slits or openings in the vent element 202.
[0102] Figure 2c A variant of the in-ear headphone device 102 as shown according to another embodiment is shown. Figure 2aAnother variant of the in-ear headphone device 102 shown. In this embodiment, in addition to the microphone 103, the in-ear headphone device further includes a feedback microphone 204. The feedback microphone is shown disposed adjacent to the ear canal-facing portion of the in-ear headphone device 102. However, according to other embodiments, the feedback microphone 204 may also be disposed towards the center of the interior of the in-ear headphone device 102 and may be acoustically coupled to the ear canal 109 via a microphone duct (not shown in the figure). The feedback microphone 204 is arranged to pick up the acoustic sounds in the ear canal 109 and feed the recorded signal to the signal processor 104. Specifically, the feedback microphone may detect the sound pressure level over the entire frequency range, which at least includes low frequencies (such as frequencies in the range of 50 Hz to 1 kHz (an example of the vowel-dominated frequency range)) and higher frequencies (such as frequencies in the range of 2 kHz to 4 kHz (an example of the consonant-dominated frequency range)). Generally, such a microphone will be configured to detect at least the entire frequency range audible to humans (i.e., the hearing range), which is typically frequencies in the range of 20 Hz to 20 kHz.
[0103] Figure 2d Another embodiment is shown, which is a variant of the in-ear headphone device 102 as Figure 2c shown. As shown, the in-ear headphone device 102 includes a damping vent 107, which includes a venting element 202 and a damping element 203, similar to the damping vent 107 described with respect to Figure 2b In another embodiment, the damping characteristic of the damping vent 107 is provided by a damping cloth at both ends of the damping vent 107, and in other embodiments, the damping characteristic of the damping vent 107 is provided by slits or openings in the venting element 202.
[0104] Figures 3a to 3h Various layouts of the vent 107 for the acoustic path of the in-ear headphone device 102 suitable for embodiments according to the present invention are shown. It should be noted that throughout the figures, damping vents are shown. However, according to other embodiments of the present invention, all the shown vents may also be used without damping elements.
[0105] Figure 3a A side view of the damping vent 107 according to an embodiment of the present invention is shown. The damping vent 107 includes a venting element 202 in the form of a cylinder and a damping element 203 in the form of a damping cloth. Although the venting element 202 is shown as a cylindrical element in this embodiment, other geometries are also conceivable.
[0106] The damping element 203 in the form of a damping cloth is shown at one end of the ventilation element 202, however, it can be located at any end of the ventilation element 202, and in another embodiment of the present invention, the damping ventilation port 107 includes damping elements 203 at both ends of the damping ventilation port 107. The damping element 203 of this embodiment is positioned within the opening of the ventilation element 202, however, in another embodiment of the present invention, the damping element 203 can be positioned in a manner that covers the opening of the ventilation element 202.
[0107] Figure 3b A side view of the damping ventilation port 107 according to an embodiment of the present invention is shown. A plurality of ventilation elements 202 form a branched damping ventilation port 107, and the branched damping ventilation port 107 further includes a damping element 203 in the form of a damping cloth. The damping element 203 of this embodiment is positioned within the opening of the ventilation element 202, however, in another embodiment of the present invention, the damping element 203 can be positioned in a manner that covers the opening of the ventilation element 202. Additionally, in other embodiments of the present invention, the branched damping ventilation port can include any number of damping elements 203, such as a damping element 203 that covers all the openings of the ventilation element 202.
[0108] Figures 3c to 3d Two side views of the damping ventilation port 107 according to an embodiment of the present invention are shown. Figure 5 c shows the damping ventilation port 107 constructed together with the speaker duct 106, and the speaker 105 can be acoustically coupled to the speaker duct 106. In this embodiment of the present invention, the speaker duct 106 and the damping ventilation port 107 form a cylindrical acoustic tube, that is, each of the two has a semi-cylindrical geometry. In other embodiments of the present invention, the speaker duct 106 and the damping ventilation port 107 can form a combined acoustic tube with any geometric shape combination. In Figure 3c it, a dashed line c-c representing the plane c is shown. In Figure 3e it, a view of the embodiment as seen from the plane c is shown, showing the longitudinal geometry of the combined speaker duct 106 and damping ventilation port 107.
[0109] Figure 3d An embodiment of the present invention is shown, in which the in-ear headphone device 102 (not shown in the figure) includes two separate damping ventilation ports 107. Each damping ventilation port 107 is similar to the damping ventilation port 107 shown in the embodiment regarding Figure 3a . Similarly, Figure 3d the configuration of the damping ventilation port 107 in it includes a ventilation element 202 and a damping element 203. The damping element 203 of this embodiment is a damping cloth present in the opening of the ventilation element 202, however, other configurations of the damping element are also conceivable.
[0110] Figure 3f An embodiment of the present invention is shown, in which the damping characteristic of the damping vent 107 is achieved by a damping element 203 in the form of a slit. In another embodiment, the damping element 203 is integrated into the venting element 202, for example to disrupt the air flow or to achieve air leakage.
[0111] Figure 3g An embodiment of the present invention is shown, in which a microphone (such as the feedback microphone 204) is arranged to primarily record sound from the damping vent 107. Thus, the microphone can be considered acoustically coupled to the venting element 202 of the damping vent 107 within the in-ear headphone device 102. In other embodiments, the in-ear headphone device 102 includes a number of venting elements 202, and according to an embodiment of the present invention, the microphone and / or the speaker can be coupled to any one of these venting elements 202. In Figure 3g the embodiment shown, the damping vent 107 has a single damping element 203 on one side. In such an embodiment, the microphone can thus primarily record sound from the external environment or primarily record sound from the ear canal, depending on the precise positioning of the damping element 203 and the microphone.
[0112] Figure 3h An embodiment of the present invention is shown, in which the speaker duct 106 and the damping vent are partially coupled by a damping element 203. The damping vent 107 also includes damping elements 203 at both ends of the venting element 202. The speaker duct 106 and the damping vent 107 can be characterized by any type of separator according to an embodiment of the present invention. The speaker 105 can be acoustically coupled to the damping vent 107 within the in-ear headphone device 102, acoustically decoupled from the damping vent 107 within the in-ear headphone device 102 (see for example Figure 3c ), or partially coupled to the damping vent 107 within the in-ear headphone device 102, as Figure 3h shown.
[0113] In the above embodiments of the present invention, various configurations of the damping vent 107 are shown. However, the present invention is not limited to any specific configuration, and thus those skilled in the art can obtain various other embodiments. The damping vent configuration can be achieved by any combination of the above embodiments; thus, the damping vent configuration can include one or more damping vents 107, each damping vent can include any number of venting elements 202 and damping elements 203, the microphone and / or the speaker can be acoustically coupled to the venting element or can have a separate duct, and the vents and ducts can have any geometry. In addition, as already mentioned, according to other embodiments of the present invention, all the vents shown in Figures 3a to 3h can be used without the damping element 203.
[0114] Figure 4 shows the characteristics of the acoustic path 501 and the electro-acoustic path 502 according to an embodiment of the present invention. The figure shows a horizontal axis representing the frequency (f) in Hertz (Hz). As shown, the frequency axis includes two frequency ranges, the vowel-dominated frequency range VDF and the consonant-dominated frequency range CDF. The vowel-dominated frequency range VDF includes frequencies in the range of 50 Hz to 800 Hz, and the consonant-dominated frequency range includes frequencies in the range of 2000 Hz (2 kHz) to 4000 Hz (4 kHz). Although the two frequency ranges are shown as two distinct ranges, this does not exclude the possibility that signal content related to vowels may exist outside the vowel-dominated frequency range VDF, and signal content related to consonants may exist outside the consonant-dominated frequency range CDF. In this context, the consonant-dominated frequency range CDF is considered to include frequencies higher than those contained in the vowel-dominated frequency range VDF. In the presence of party noise or the like, most of the noise energy falls within the vowel-dominated frequency range VDF. The figure also shows the passbands of the acoustic path 501 and the electro-acoustic path 502 of the in-ear headphone device 102. The acoustic path 501 includes at least the vent 107 (see, for example Figure 1 ), and the electro-acoustic path 502 includes at least the microphone 103, the signal processor 104, and the speaker 105. The acoustic path 501 can be any of the acoustic paths previously described, and the electro-acoustic path 502 can be any of the electro-acoustic paths previously described. As shown, the acoustic path 501 is concentrated on the vowel-dominated frequency range VDF. The acoustic path 501 is effectively a low-pass filter, where the vowel-dominated frequency range VDF is within the passband of the acoustic path 501. However, the electro-acoustic path 502 processes a much wider frequency range than the acoustic path 501 and covers both the vowel-dominated frequency range VDF and the consonant-dominated frequency range CDF. Figure 4 Also shown is a vertical arrow extending from the electro-acoustic path 502 to the acoustic path 501 within the vowel-dominated frequency range VDF. The arrow represents the compensation performed by the electro-acoustic path 502. This compensation can be best understood by considering Figure 5 .
[0115] Figure 5shows the concept of using an electroacoustic path 502 to compensate for the contribution from the acoustic path 501 in the vowel-dominated frequency range VDF. Speech includes both vowels and consonants, and speech intelligibility is largely attributed to the correct detection of consonants. However, in many situations, such as in a cocktail party scenario, the vowels from competing speakers carry a large amount of the acoustic energy of the speech and have a masking effect on the consonants of the conversation participants. In other words, the signal content in the vowel-dominated frequency range VDF can impose a masking effect on the signal content present in the consonant-dominated frequency range CDF. For this reason, the in-ear headphone device of the speech intelligibility enhancement system (see, for example Figure 1 and Figures 2a to 2d the in-ear headphone device) is arranged such that by using the electroacoustic path 502 to compensate for the contribution from the acoustic path 501 in the vowel-dominated frequency range VDF, the difference 503 between the resulting sound pressure level in the ear canal 109 contributed by the vowel-dominated frequency range VDF and the resulting sound pressure level in the ear canal 109 contributed by the consonant-dominated frequency range CDF is reduced. Figure 5 shows the resulting sound pressure level (SPL) 506 contributed by the vowel-dominated frequency range VDF and the resulting sound pressure level 507 contributed by the consonant-dominated frequency range. As seen in this embodiment, the resulting sound pressure level 506 contributed by the vowel-dominated frequency range is present at the center frequency 504 of the vowel-dominated frequency range VDF, and the resulting sound pressure level 507 contributed by the consonant-dominated frequency range is present at the center frequency 505 of the consonant-dominated frequency range CDF. However, in other embodiments, the resulting sound pressure level may represent the average sound pressure level of the entire vowel-dominated frequency range and consonant-dominated frequency range, or the average sound pressure level of a sub-range thereof. In Figure 5A difference 503 between two resulting sound pressure levels is seen. In a preferred embodiment, the difference 503 is kept below 15 dB, and in a more preferred embodiment, the difference 503 is kept below 10 dB. This keeping may require reducing the difference 503, and this is achieved by compensating for the contribution from the acoustic path 501 in the vowel-dominated frequency range VDF using the electroacoustic path 502. According to embodiments of the present invention, this compensation can be achieved in a variety of ways. In the present embodiment, the signal processor 104 of the voice intelligibility enhancement system 101 is substantially arranged to apply a phase shift to the signal recorded by the microphone 103, so as to acoustically reproduce the phase-shifted audio signal in the vowel-dominated frequency range VDF using the speaker 105. Importantly, the phase-shifted audio signal has an opposite effect on the acoustic sound in the ear canal 109 contributed by the acoustic path 501, and this effect ensures that the overall transfer function of the sound from the external acoustic environment 108 to the ear canal 109 exhibits the characteristic that the difference 503 is below a specified level. More importantly, the compensation is not intended to completely counteract the acoustic sound in the ear canal 109 contributed by the acoustic path 501, because it is still the goal to achieve a certain degree of natural reproduction of the acoustic sound in the passband of the acoustic path 501. This is particularly important because the vowels produced by the conversation partner are themselves speech cues and also establish the time window in which key consonants may appear. Obviously, reducing the difference 503 as described above is a way to improve the signal-to-mask ratio.
[0116] In a preferred embodiment, the signal processing algorithm adjusts the amount of attenuation applied to the vowel-dominated frequency range VDF according to the sound pressure level, such that when the low-frequency level is low enough that consonant masking is unlikely to occur, the sound is perceived with a natural and / or desired spectral balance.
[0117] In another preferred embodiment, the signal processing algorithm is arranged to detect when the wearer of the in-ear headphone device is speaking and adjust the compensation so as to maintain a natural impression of the wearer's own speech.
[0118] Figures 6 to 15 is shown in relation to the spectrum associated with the in-ear headphone device 102 according to an embodiment of the present invention as Figure 16 shown. In the present embodiment, the in-ear headphone device 102 of the voice intelligibility enhancement system 101 includes two microphones. Further details regarding this embodiment are given in the accompanying Figure 16 text.
[0119] Figure 6Shows the spectra of three signals that occur in the absence of any baffle effect, i.e., as if the signals were recorded using an omnidirectional microphone located at the position where the wearer of the speech intelligibility enhancement system 101 is standing. The figure shows three signal curves S1, S2, and S3 plotted on a graph, which shows the amplitude spectral density (ASD) (in dB re 20 micro-Pascals per square root of Hertz [dB re 20 μPa / sqrt(Hz)]) as a function of frequency (in Hertz [Hz]).
[0120] The signal curve S1 represents the long-term average spectrum of the noise present in a room with thirty people speaking. Throughout the following description, this will be referred to as babble noise.
[0121] The signal curve S2 also represents the long-term average spectrum of a single speaker located approximately one meter away from the wearer of the speech intelligibility enhancement system 101. Any pauses in speech made by the speaker have been omitted from the integration, resulting in the signal curve S2.
[0122] The signal curve S3 represents the short-term spectrum of the consonant "t" spoken by a single speaker located one meter away from the wearer of the speech intelligibility enhancement system 101. As shown, the spectrum of the consonant "t" peaks at approximately 3 kHz, i.e., within the consonant dominant frequency range CDF. It should be noted that the consonant "t" was chosen solely for demonstration purposes, and those skilled in the art will be aware of the spectra of other consonants, which can be readily substituted and demonstrate the same principles as will be elaborated below.
[0123] The three signal curves S1, S2, and S3 together demonstrate a typical cocktail party situation where speech is difficult to understand due to the presence of babble noise, which has a masking effect on consonants that is crucial for speech intelligibility. Hereinafter, the three signal curves S1, S2, and S3 can be regarded as input signals for signal processing by the acoustic path and electroacoustic path of at least one in-ear headphone device. This signal processing is illustrated by Figures 7 to 9 the transfer function.
[0124] Figure 7Four simplified transfer functions T1 (square), T2 (cross), T3 (triangle), and T4 (circle) are shown. The transfer functions are simplified in the sense that they do not account for ear canal resonances. The transfer functions are plotted on a graph that shows the real-ear gain (REG) (in decibels (dB)) according to frequency (in Hz). Transfer function T1 is the transfer function of the vent 107 of the acoustic path. Hereinafter, this transfer function is referred to as the vent transfer function. Transfer function T2 is the transfer function of the audio signal recorded by one of the multiple microphones in the in-ear headphone device, and in this case, this microphone acts as a pressure microphone, such as an omnidirectional microphone. Hereinafter, this transfer function is referred to as the omnidirectional microphone transfer function. Transfer function T3 is the transfer function of the audio signal recorded by one or more microphones of the in-ear headphone device, which acts as a supercardioid directional microphone. Since the directional microphone is most sensitive in a specific direction, it is less sensitive to acoustic sounds with more diffuse characteristics such as multi-coincidence noise. This is reflected by transfer function T4, which is the transfer function of the directional microphone when subjected to diffuse acoustic sounds (e.g., multi-coincidence noise). As can be seen by comparing transfer functions T3 and T4, the directional microphone effectively suppresses diffuse acoustic sounds by approximately 5.5 dB compared to acoustic sounds with directional characteristics. This shows that the directional microphone is more sensitive to a speaker standing in front of the wearer of the speech intelligibility enhancement system 101 than to the multi-coincidence noise present in the room.
[0125] Figure 8 and Figure 9 shows the corresponding phase diagrams and delay diagrams of the transfer functions as Figure 7 shown. In Figure 8 four phase curves P1 (square), P2 (cross), P3 (triangle), and P4 (circle) are shown, which respectively correspond to the four transfer functions T1, T2, T3, and T4. Figure 8 The curves in Figure 9 show the phase (in degrees) according to frequency (in Hz). In Figure 9 four group delay curves D1 (square), D2 (cross), D3 (triangle), and D4 (circle) are shown, which respectively correspond to the four transfer functions T1, T2, T3, and T4. Figure 9 The curves in
[0126] show the delay (in microseconds) according to frequency (in Hz). To illustrate that the present invention according to the present embodiment can arrange the processing latency, a fixed delay of 100 microseconds has been added to the transfer function of the electroacoustic path.
[0126] During the identification of the types of acoustic sound signals present in the room during a cocktail party situation (see Figure 6)and the transfer function of at least one in-ear headphone device 102 of the voice intelligibility enhancement system (see Figure 7 )After that, the effects of applying the transfer function to these signals are discussed with reference to the following figures.
[0127] Figure 10 Shows the effect of applying the vent transfer function T1 to three input audio signals represented by signal curves S1, S2, and S3. Figure 10 The graph on shows in the same way as Figure 6 The graph on shows the amplitude spectral density (ASD) according to frequency. The signal curve S4 shows the result of applying the vent transfer function T1 to the acoustic sound signal represented by the signal curve S1. In other words, the signal curve S4 shows the contribution of the vent / acoustic path to the multiple coincident noises present in the ear canal of the wearer of the in-ear headphone device 102. The signal curve S5 shows the result of applying the vent transfer function T1 to the acoustic sound signal represented by the signal curve S2. In other words, the signal curve S5 shows the contribution of the vent to the sound of a specific person speaking in the ear canal of the wearer of the in-ear headphone device. The signal curve S6 shows the result of applying the transfer function T1 to the acoustic sound signal represented by the signal curve S3. In other words, the signal curve S6 shows the contribution of the vent to the consonants generated by a specific person speaking present in the ear canal of the wearer of the in-ear headphone device. As Figure 10 shown, the effect of the vent transfer function T1 is that the consonant (in this case the consonant "t") is suppressed relative to the lower frequency content when passing through the vent.
[0128] Figure 11 Shows the effect of applying the transfer function T3 to three input audio signals represented by signal curves S1, S2, and S3. Figure 11 The graph on shows in the same way as Figure 6 The graph on shows the amplitude spectral density (ASD) according to frequency. The signal curve S7 shows the result of applying the transfer function T3 to the acoustic sound signal represented by the signal curve S1. In other words, the signal curve S7 shows the contribution of the directional microphone to the multiple coincident noises present in the ear canal of the wearer of the in-ear headphone device 102. The signal curve S8 shows the result of applying the transfer function T3 to the acoustic sound signal represented by the signal curve S2. In other words, the signal curve S8 shows the contribution of the directional microphone to the desired voice signal present in the ear canal of the wearer of the in-ear headphone device. The signal curve S9 shows the result of applying the transfer function T3 to the acoustic sound signal represented by the signal curve S3. In other words, the signal curve S9 shows the contribution of the directional microphone to the consonant "t" present in the ear canal of the user wearing the in-ear headphone device.
[0129] Figure 12 also shows a graph of the amplitude spectral density (ASD) according to frequency in the same manner as the graph on Figure 6 . This graph shows three signal curves S10, S11, and S12. The signal curve S10 corresponds to the long-term average spectrum of the multiplexed coincidence noise equal to the signal curve S1 in Figure 6 . The signal curve S11 corresponds to the signal curve S4 as seen in Figure 9 , that is, the signal curve 11 shows the effect of applying the vent transfer function to the multiplexed coincidence noise. When comparing the signal curves S10 and S11, the effect of the vent of the acoustic path is clearly seen, especially the inherent low-pass characteristic of the vent. The multiplexed noise with most of the energy at low frequencies (i.e., in the passband of the vent) passes through the vent, while the higher frequencies are significantly attenuated due to the presence of the vent. It can also be seen that for frequencies below 600 Hz, the presence of the vent does not significantly reduce the amplitude spectral density.
[0130] However, the signal curve S12 shows the effect produced when the vent transfer function T1 and the omnidirectional microphone transfer function T2 (see Figure 7 ) are applied to the multiplexed coincidence noise signal S10 and combined in the ear canal. Essentially, the signal curve S12 represents the long-term average spectrum of the multiplexed coincidence noise present in the ear canal of the user of the in-ear headphone device 102. As shown, the multiplexed coincidence noise is significantly reduced compared to the multiplexed coincidence noise in the case where the in-ear headphone device is not present (see the signal curve S10). It can be seen from these example transfer functions that a reduction of approximately 6 dB is achieved at 300 Hz.
[0131] Figure 13 also shows a graph of the amplitude spectral density (ASD) according to frequency in the same manner as the graph on Figure 6 . Specifically, the graph shows three signal curves S13, S14, and S15. The signal curve S13 corresponds to the long-term average spectrum of a single speaker approximately one meter away from the wearer of the speech intelligibility enhancement system 101, that is, the signal curve S13 corresponds to the signal curve S2 as shown in Figure 6 . The signal curve S14 shows the effect of applying the vent transfer function T1 to the long-term average spectrum of the speaker's speech. Therefore, the signal curve S14 directly corresponds to the signal curve S5 in Figure 10 . The signal curve S15 represents the resulting speech signal present in the ear canal of the user wearing the in-ear headphone device 102, thus representing the contributions of the acoustic path and the electroacoustic path of the in-ear headphone device.
[0132] Figure 14 also shows a graph of the amplitude spectral density (ASD) according to frequency in the same manner as the graph on Figure 6The graph of the amplitude spectral density (ASD) according to frequency is shown in the same manner as the graph on Figure 6 . Specifically, the graph shows three signal curves S16, S17, and S18. The signal curve S16 corresponds to Figure 14 the signal curve S3 in
[0133] Figure 15 and thus represents the short-time average spectrum of the consonant "t" produced by a speaker standing approximately 1 meter away from the wearer of the in-ear headphone device 102. The signal curve S17 shows the effect of applying the vent transfer function T1 to the short-time average spectrum of the consonant "t", and as Figure 6 shown, the vent significantly attenuates the signal. This is not surprising when looking at the vent transfer function T1 with low-pass characteristics. The signal curve S18 shows the resulting consonant signal present in the ear canal of a user wearing the in-ear headphone device, thus representing the contributions of the acoustic path and the electroacoustic path of the in-ear headphone device. In this example, amplification of the consonant is achieved (which is obvious when comparing the signal curve S18 with the signal curve S16). The advantage of such amplification is that it further improves speech intelligibility, as is obvious from the following figure. Figure 6 The graph of the amplitude spectral density (ASD) according to frequency is also shown in the same manner as the graph on Figure 6 . Specifically, the graph shows four signal curves S19, S20, S21, and S22. The signal curve S19 corresponds to the multiple coincidence noise signal also seen as the signal curve S1 in Figure 12 , and the signal curve S20 corresponds to the consonant signal also seen as the signal curve S3 in Figure 14 . When directly comparing the signal curves S19 and S20, it can be seen that if the user is not wearing the in-ear headphone device of the speech intelligibility enhancement system, the low-frequency multiple coincidence noise is present at a high level compared to the consonant "t". In this example, the multiple coincidence noise exerts a masking effect on the consonant "t". Note that this example only involves the letter "t", but similar (and even more obvious) effects generally exist for other consonants. This makes speech understanding particularly difficult because the consonants produced by the speaker of interest are "drowned out" by the multiple compound noise present from other people in the room. The signal curve S21 corresponds to the signal curve S12 as seen in Figure 15When directly comparing, the beneficial effects can be truly understood first. As shown in the figure, in the low-frequency range of the spectrum, that is, in the vowel-dominated frequency range, the signal curve 21 decreases compared to the signal curve S19. Effectively, this indicates that the electroacoustic path is arranged (through a specific transfer function) such that it compensates for the contributions from the acoustic path / vents in the vowel-dominated frequency range, so that these contributions have a lower masking effect on consonants in the consonant-dominated frequency range. Thus, improved speech intelligibility is achieved. Therefore, Figure 15 It shows that the contribution from the acoustic path in the vowel-dominated frequency range is compensated by the electroacoustic path to improve the signal-to-masking ratio.
[0134] Figure 16 is also in the same way as Figure 6 The graph showing the amplitude spectral density (ASD) according to frequency is shown in the same way as the graph above. Specifically, this graph shows four signal curves S23, S24, S25, and S26. The signal curve 23 corresponds to the signal curve S1 (see Figure 6 ), the signal curve S24 corresponds to the signal curve S2 (also see Figure 6 ), the signal curve S25 corresponds to the signal curve S12 (see Figure 12 ), and the signal curve S26 corresponds to the signal curve S15 (see Figure 13 ). This also reveals the beneficial effects on the long-term average spectrum of speech. If the in-ear headphone device 102 of the speech intelligibility enhancement system 101 is not used, at high frequencies, such as frequencies in the range of 1 kHz to 6 kHz, the multiplexed coincidence noise spectrum is higher than the long-term average spectrum of speech. However, when inserted into the ear canal, the speech intelligibility enhancement system 101 improves the speech-to-masking energy ratio, especially in the high-frequency range. This has the following benefits: Speech cues from the speaker of interest can be more easily detected by the user of the system, thus positively affecting speech intelligibility.
[0135] Figure 17 It shows the in-ear headphone device 102 of the speech intelligibility enhancement system according to an embodiment of the present invention. The in-ear headphone device 102 is arranged to apply transfer functions T1 to T4 as Figure 7 shown. Therefore, by using the in-ear headphone device 102, all the results of signal processing as Figures 6 to 16 shown can be achieved. The in-ear headphone device of this embodiment includes two microphones 103 arranged to record the acoustic sounds present in the external environment. The microphones of this embodiment are two omnidirectional microphones that are combined using a signal processor 104 to achieve the desired directivity characteristics. However, dedicated directional microphones can also be used according to another embodiment of the present invention.
[0136] It should be noted that the speech intelligibility enhancement system 101, as mentioned in any of the foregoing descriptions, may include two in-ear headphone devices 102; one in-ear headphone device 102 is used for each ear of the wearer of the speech intelligibility enhancement system.
[0137] List of reference numerals:
[0138] 101 Speech intelligibility enhancement system
[0139] 102 In-ear headphone device
[0140] 103 Microphone
[0141] 104 Signal processor
[0142] 105 Speaker
[0143] 106 Speaker duct
[0144] 107 Vent
[0145] 108 External acoustic environment
[0146] 109 Ear canal
[0147] 110 Auricle (outer ear)
[0148] 111 Flexible earplug
[0149] 201 Tympanic membrane (eardrum)
[0150] 202 Venting element
[0151] 203 Damping element
[0152] 204 Feedback microphone
[0153] 501 Acoustic path
[0154] 502 Electroacoustic path
[0155] 503 Difference in sound pressure level
[0156] 504 Center frequency of the vowel-dominated frequency range
[0157] 505 Center frequency of the consonant-dominated frequency range
[0158] 506 Resultant sound pressure level contributed by VDF
[0159] 507 Resultant sound pressure level contributed by CDF
[0160] VDF Vowel-dominated frequency range
[0161] CDF Consonant-dominated frequency range
[0162] S1 - S26 Signal Curves
[0163] T1 - T4 Transfer Functions
[0164] P1 - P4 Phase Curves
[0165] D1 - D4 Delay Curves.
Claims
1. A speech intelligibility enhancement system for difficult acoustic conditions, the speech intelligibility enhancement system comprising at least one in-ear headphone device for insertion into a person's ear canal, the at least one in-ear headphone device being arranged with an ear canal-facing portion and an environment-facing portion, and the at least one in-ear headphone device comprises: an acoustic path including a vent, the acoustic path coupling the environment-facing portion to the ear canal-facing portion; and an electroacoustic path including a microphone at the environment-facing portion, a filter, and a speaker at the ear canal-facing portion; wherein the acoustic path is arranged to transmit acoustic sounds in a vowel-dominated frequency range, and wherein the electroacoustic path is arranged to acoustically reproduce sound signals in a consonant-dominated frequency range and in the vowel-dominated frequency range; and wherein the electroacoustic path is arranged such that in the vowel-dominated frequency range, the contribution from the acoustic path is compensated through the electroacoustic path to improve the signal-to-mask ratio.
2. The speech intelligibility enhancement system according to claim 1, wherein, improving the signal-to-mask ratio includes increasing the resulting sound pressure level present in the consonant-dominated frequency range relative to the resulting sound pressure level present in the vowel-dominated frequency range.
3. The speech intelligibility enhancement system according to claim 1 or 2, wherein, improving the signal-to-mask ratio includes: reducing the difference between the resulting sound pressure level in the ear canal contributed by the vowel-dominated frequency range and the resulting sound pressure level in the ear canal contributed by the consonant-dominated frequency range by using the electroacoustic path to compensate for the contribution from the acoustic path in the vowel-dominated frequency range.
4. The speech intelligibility enhancement system according to any one of the preceding claims, wherein, the consonant-dominated frequency range includes frequencies higher than the vowel-dominated frequency range.
5. The speech intelligibility enhancement system according to any one of the preceding claims, wherein, the acoustic path is arranged with an acoustic transfer function having a low-pass characteristic, the low-pass characteristic having a passband and a cut-off frequency, and wherein the vowel-dominated frequency range includes frequencies lower than the cut-off frequency, and wherein the consonant-dominated frequency range includes frequencies higher than the cut-off frequency.
6. The speech intelligibility enhancement system according to any one of the preceding claims, wherein, the cut-off frequency is in the range of 250 Hz to 4 kHz.
7. The speech intelligibility enhancement system according to any one of the preceding claims, wherein, the vowel-dominated frequency range includes frequencies in the range of 50 Hz to 1 kHz.
8. The speech intelligibility enhancement system according to any one of the preceding claims, wherein, the consonant-dominated frequency range includes frequencies in the range of 2 kHz to 4 kHz.
9. The speech intelligibility enhancement system according to any one of the preceding claims, wherein, the difference is below 15 dB, such as below 10 dB, such as below 8 dB, such as below 6 dB, for example below 5 dB.
10. The voice intelligibility system according to any one of the preceding claims, wherein, the electroacoustic path is arranged to compensate for the contribution from the acoustic path within a signal processing frequency range of 300 Hz to 1 kHz.
11. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the compensation of the contribution from the acoustic path is signal-dependent.
12. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the compensation of the contribution from the acoustic path is level-dependent.
13. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the electroacoustic path is arranged to compensate for the contribution from the acoustic path by reproducing the sound signal in at least a part of the vowel-dominated frequency range.
14. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the electroacoustic path is arranged to reproduce the sound signal in at least a part of the vowel-dominated frequency range with a polarity opposite to that of the acoustic sound transmitted by the acoustic path, and the effect of the compensation is that, compared with the case where at least one in-ear headphone device is not inserted into the user's / wearer's ear canal, the perceived loudness of the acoustic sound in the vowel-dominated frequency range is reduced.
15. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the electroacoustic path is arranged to reproduce the sound signal in at least a part of the vowel-dominated frequency range by applying a phase shift to the sound signal.
16. The voice intelligibility enhancement system according to claim 15, wherein, the phase shift is greater than 90 degrees and less than 270 degrees.
17. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the vent is a damped vent.
18. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the speaker and the vent are acoustically separated inside the at least one in-ear headphone device.
19. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the vent is arranged to have a cross-sectional area equal to that of a cylinder, and the diameter of the cylinder is in the range of 1.5 mm to 3.5 mm, such as 2.0 mm to 3.0 mm, for example 2.3 mm or 2.5 mm.
20. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the vent is arranged to have a length equal to that of a cylinder, and the length of the cylinder is in the range of 2.5 mm to 10 mm, such as 3.5 mm to 9 mm, such as 4.5 mm to 8 mm, for example 5 mm or 7 mm.
21. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the filter is arranged in the signal processor of the at least one in-ear headphone device, such as a digital signal processor.
22. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the at least one in-ear headphone device is battery-powered, such as powered by a rechargeable battery.
23. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the at least one in-ear headphone device comprises two in-ear headphone devices, each ear canal of the person uses one in-ear headphone device, and wherein the two in-ear headphone devices are arranged to coordinate their settings with each other.
24. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the at least one in-ear headphone device comprises a feedback microphone at the ear canal facing portion.
25. The voice intelligibility enhancement system according to claim 24, wherein, the electroacoustic path is arranged to compensate for the contribution from the acoustic path based on the input provided by the feedback microphone.
26. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the microphone of the electroacoustic path is a directional microphone.
27. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the electroacoustic path of the at least one in-ear headphone device comprises a plurality of microphones.
28. The voice intelligibility enhancement system according to any one of the preceding claims, wherein, the electroacoustic path is arranged to amplify sound with a nominal gain in the passband of the electroacoustic path.
29. A method for enhancing voice intelligibility under difficult acoustic conditions, the method comprising the steps of: inserting at least one in-ear headphone device into the ear canal of a person, the at least one in-ear headphone device being arranged with an ear canal facing portion and an environment facing portion, the at least one in-ear headphone device comprising an acoustic path and an electroacoustic path, the acoustic path comprising a vent hole coupling the environment facing portion and the ear canal facing portion, the electroacoustic path comprising a microphone, a filter at the environment facing portion and a speaker at the ear canal facing portion; transmitting acoustic sounds in the vowel-dominated frequency range from the environment facing portion to the ear canal facing portion through the acoustic path; acoustically reproducing sound signals in the consonant-dominated frequency range and the vowel-dominated frequency range through the electroacoustic path; and compensating for the contribution from the acoustic path in the vowel-dominated frequency range through the electroacoustic path, so that the signal-to-mask ratio is improved.
30. The method according to claim 29, wherein, the method is executed by the voice intelligibility enhancement device according to any one of claims 1-28.