Hearing system comprising a hearing instrument and method for operating a hearing instrument

By analyzing the user's own voice reference signal, determining the stress rhythm pattern, and enhancing or superimposing matching speech stress in the hearing instrument, the problem of insufficient speech perception in noisy environments is solved, and a better speech perception effect is achieved.

CN115706910BActive Publication Date: 2025-11-11SIVANTOS PTE LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210949691.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-08-16
Filing Date
2022-08-09
Publication Date
2025-11-11
Estimated Expiration
2042-08-09

AI Technical Summary

Technical Problem

Existing hearing aids are inadequate in speech perception when compensating for hearing loss, especially in noisy environments, and commonly used signal processing methods result in distortion and are not effective enough.

Method used

By analyzing the user's self-voice reference signal, the stress rhythm pattern (SRP) is determined, and during the speech recognition process, the speech stress that matches the user's SRP is enhanced, or artificial speech stress is superimposed to adapt to the user's SRP, thereby improving speech perception.

Benefits of technology

It significantly improves the ability of hearing instrument users to perceive speech in noisy environments, enhances the comprehensibility of speech stress, and adapts to users' personalized needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115706910B_ABST
    Figure CN115706910B_ABST
Patent Text Reader

Abstract

This invention relates to a hearing system (2) including a hearing instrument (4) and a method for operating the hearing instrument (4). Here, sound signals from the environment of the hearing instrument (4) are captured by the hearing instrument (4), the captured sound signals are processed, and the processed sound signals are output to the user of the hearing instrument (4). In the speech recognition step (31), the captured sound signals are analyzed to identify speech intervals, wherein the captured sound signals contain speech. In the speech enhancement process (36) performed during the identified speech intervals, the amplitude of the processed sound signals is periodically changed according to a time pattern consistent with the user's repetition rhythm pattern.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for operating a hearing instrument. It also relates to a hearing system including the hearing instrument. Background Technology

[0002] Typically, a hearing aid is an electronic device designed to support the hearing of the person wearing it (referred to as the user or wearer of the hearing aid). In particular, this invention relates to a hearing aid specifically configured to at least partially compensate for hearing loss in a user with hearing impairment. Other types of hearing aids are designed to support the hearing of users with normal hearing, i.e., to improve speech perception in complex acoustic conditions.

[0003] Hearing aids are typically designed to be worn in or on a user's ear, for example, as behind-the-ear (BTE) or in-the-ear (ITE) devices. Internally, hearing aids generally include an (electroacoustic) input transducer, a signal processor, and an output transducer. During operation, the input transducer captures sound signals from the environment and converts them into input audio signals (i.e., electrical signals that transmit sound information). In the signal processor, the captured sound signals (i.e., the input audio signals) are processed, particularly amplified according to sound frequencies, to support the user's hearing, especially to compensate for hearing impairment. The signal processor outputs the processed audio signal (also called the processed sound signal) to the output transducer. Typically, the output transducer is an electroacoustic transducer (also called a "receiver") that converts the processed sound signal into processed airborne sound, which is emitted into the user's ear canal. Alternatively, the output transducer can be an electromechanical transducer that converts the processed sound signal into structurally propagating sound (vibrations), transmitted to, for example, the user's skull. In addition to the classic hearing devices mentioned above, there are also implantable hearing devices, such as cochlear implants, and hearing devices whose output converters output processed sound signals by directly stimulating the user's auditory nerve.

[0004] The term "hearing system" refers to a device or device component and / or other structure that provides the functionality required for the operation of a hearing instrument. A hearing system may consist of a single hearing instrument. Alternatively, a hearing system may include a hearing instrument and at least one additional electronic device, which may be, for example, another hearing instrument for the user's other ear, a remote control, and a programming tool for the hearing instrument. Furthermore, modern hearing systems typically include a hearing instrument and software applications for controlling and / or programming the hearing instrument, which are installed on or can be installed on a computer or a mobile communication device such as a mobile phone (smartphone). In the latter case, the computer or mobile communication device is usually not part of the hearing system. In particular, most frequently, the computer or mobile communication device is manufactured and sold independently of the hearing system.

[0005] A typical problem for people with hearing loss is poor speech perception, which is often caused by inner ear pathology that leads to a reduced dynamic range in individuals with hearing loss. This means that soft sounds become inaudible to hearing-impaired listeners (especially in noisy environments), while loud sounds more or less maintain their perceived loudness level.

[0006] Hearing aids typically compensate for hearing loss by amplifying the captured sound signal. Therefore, compression is often used to compensate for the reduced dynamic range of hearing-impaired users; that is, increasing the amplitude of the processed sound signal as a function of the input signal level. However, due to real-time limitations in signal processing, the compression methods commonly used in hearing aids often lead to various technical problems and distortions. Moreover, in many cases, compression is insufficient to enhance speech perception to a satisfactory degree.

[0007] To improve speech perception for users wearing hearing aids, EP 3 823 306 A1 discloses a method for operating a hearing aid, wherein speech analysis is performed on sound signals captured from the environment of the hearing aid. During speech intervals, i.e., within time intervals (in which the captured sound signals contain speech), at least one time derivative of the amplitude and / or pitch of the captured sound signal is determined. If at least one derivative satisfies a predefined criterion, such as exceeding a predefined threshold, the amplitude of the processed sound signal is temporarily increased. Known methods allow for the detection and enhancement of rhythmic speech stress (“speech accent”), i.e., changes in the amplitude and / or pitch of speech, thereby significantly improving speech perception for users of hearing aids.

[0008] A hearing instrument incorporating various speech enhancement algorithms is known from EP 1 101 390 B1. Here, the level of speech segments in an audio stream is increased. Speech segments are identified by analyzing the envelope of the signal levels. Specifically, sudden level peaks (bursts) are detected as indicators of speech. Summary of the Invention

[0009] The technical problem to be solved by the present invention is to provide a method for operating a hearing instrument, which provides users wearing hearing instruments with further improved speech perception.

[0010] Another technical problem to be solved by the present invention is to provide a hearing system including a hearing instrument that provides users wearing the hearing instrument with further improved speech perception.

[0011] According to the present invention, the above-mentioned technical problems are solved independently by the method for operating a hearing instrument according to the present invention and the hearing system according to the present invention. Preferred embodiments of the present invention are described in the following description.

[0012] According to a first aspect of the invention, a method for operating a hearing instrument designed to support the hearing of a user (particularly a user with hearing impairment) is provided. The method includes, for example, capturing sound signals from the environment of the hearing instrument via an input converter of the hearing instrument. The captured sound signals are processed, for example, by a signal processor of the hearing instrument (particularly to at least partially compensate for the user's hearing impairment), thereby generating a processed sound signal. The processed sound signal is output to the user, for example, by an output converter of the hearing instrument. In a preferred embodiment, the captured sound signal and the processed sound signal are audio signals, i.e., electrical signals that transmit sound information, before being output to the user.

[0013] The hearing device can be any of the types described above. Preferably, it is designed to be worn inside or on the user's ear, such as as a BTE hearing device (with internal or external receivers) or an ITE hearing device. Alternatively, the hearing device can be designed to be implantable. The processed sound signal can be as airborne sound, structure-borne sound, or as a signal output that directly stimulates the user's auditory nerve.

[0014] In the speech recognition step, during normal operation of the hearing aid, the captured sound signal is analyzed to identify (detect) speech intervals, wherein the captured sound signal contains speech. For example, methods known from EP 1 101 390 B1 can be used to detect speech and speech intervals. During the enhancement process, during the identified speech interval, the amplitude of the processed sound signal is periodically varied to enhance or induce speech stress in the processed sound signal. Specifically, the amplitude of the processed sound signal is temporarily increased. Here, the amplitude varies periodically according to a time pattern consistent with the user's stress rhythmic pattern (SRP). Therefore, the speech contained in the captured audio signal is suited to the user's SRP. It should be noted that, within the scope of the invention, the enhancement process can be applied to the captured sound signal at any stage of signal processing. Therefore, the enhancement process can be applied to the captured sound signal in its initial form or in a partially processed form.

[0015] A speaker's "stress rhythm pattern" (SRP) typically describes a single stressed rhythm, which is a temporal pattern of stress (in language) that includes temporary variations (peaks) in the amplitude and / or pitch of the speaker's voice, used (consciously or unconsciously) to construct and emphasize speech. Typically, stress used in speech has a rhythmic structure, meaning that the SRP of speech is repeated continuously in a different but similar manner, which is unique to the speaker. Subsequently, similar to EP 3 823306 A1, the smallest identifiable unit of the SRP, the single peak in amplitude and / or pitch used to construct and emphasize speech, is called a "speech stress." Typically, such speech stresses have a duration of approximately 5 to 15 msec (milliseconds) and occur over time intervals of more than 400 milliseconds (corresponding to a rate of less than 2.5 Hz).

[0016] This invention is based on the experience that if the speaker and listener have similar SRPs (Speech Representation Points), the listener can more easily understand the speech, while if the speaker and listener have significantly different SRPs, the speech is more difficult to understand. Utilizing this experience, this invention proposes to artificially distort the speech contained in the captured sound signal so that the speech is closer to the listener's, i.e., the user of the hearing instrument's SRP. It has been found that this can significantly improve the user's speech perception, thereby outweighing the negative effects caused by distortion of the original speech.

[0017] Within the scope of this invention, the user's SRP can be predefined or determined independently of the operation of the hearing instrument. However, in a preferred embodiment of the invention, determining the user's SRP is performed as part of the method of operating the hearing instrument. For this purpose, preferably, the method further includes an own-voice (OV) analysis process, wherein the user's SRP is determined from an own-voice reference signal (OV reference signal) containing the user's speech.

[0018] Note that the “self-voice reference signal” (“OV reference signal”) mentioned above is different from the “captured sound signal” processed during the enhancement process. While the latter is acquired during normal operation of the hearing instrument (particularly during time intervals of user silence), the OV reference signal is typically collected in steps prior to normal operation of the hearing instrument (particularly during the enhancement process). Preferably, the OV reference signal is collected by the hearing instrument during a setup step, particularly using the hearing instrument’s input transducer. However, within the scope of this invention, the OV reference signal can also be captured and optionally analyzed externally using a separate part of the hearing system. For example, software installed on a smartphone can be used to capture the OV reference signal. In a suitable embodiment of the method, the setup step is performed once during an initial fitting process, in which the hearing instrument is initially set to the needs and preferences of an individual user. Alternatively, the setup step can be provided such that it can be repeated at regular intervals or upon request by the user or healthcare professional. In another embodiment of the invention, the OV reference signal is collected during normal operation of the hearing instrument at self-voice intervals (OV intervals) during which the user speaks.

[0019] In a preferred embodiment of the invention, the enhancement process is performed according to a method known from EP 3 823 306 A1, the entire contents of which are incorporated herein by reference. In this embodiment, in a derivation step (performed during the identified speech interval and prior to the enhancement process), at least one derivative of the amplitude and / or at least one derivative of the pitch of the captured sound signal, i.e., the fundamental frequency, is determined. During the enhancement process, the at least one derivative is compared to a predefined criterion (e.g., a predefined threshold, see EP 3 823 306 A1), and satisfaction of the criterion is used as an indication of speech accent. That is, whenever at least one derivative satisfies the criterion, a speech accent is identified (detected). The identified speech accent is enhanced to be more easily perceived by the user by temporarily applying a gain and thus temporarily increasing the amplitude of the processed sound signal.

[0020] However, according to the present invention and unlike the teachings of EP 3 823 306 A1, only those identified speech accents in the captured sound signal that match the user's SRP are amplified. In other words, if the speech accent (in terms of its temporal occurrence in a series of speech accents) does not match the user's SRP, the identified speech accent is not amplified (i.e., the amplitude of the processed sound signal is not temporarily increased when the speech accent is identified). By selectively amplifying those identified speech accents that match the user's SRP (and by not amplifying the mismatched speech accents), the speech rhythm contained in the captured sound signal adapts to the user's SRP, thereby improving the user's speech perception.

[0021] In another embodiment of the invention, alternatively or in addition to simply enhancing the matching speech accents, at least one artificial speech accent is superimposed on the captured audio signal (in its initial or partially processed form), wherein the timing of the artificial speech accent is selected such that it matches the user's SRP in a series of previous (artificial or natural) speech accents. Here, "artificial speech accent" is a temporary increase in the amplitude of the processed audio signal superimposed on the captured audio signal, regardless of whether a natural speech accent (i.e., speech accent in natural speech) is identified in the captured audio signal at that point in time. In a variation of this embodiment, one or more artificial speech accents are created to fill gaps in a series of natural speech accents that match the user's SRP. In another variation of this embodiment, a series of artificial speech accents corresponding to the user's SRP are superimposed on the captured audio signal, regardless of natural speech accents. In other words, the captured audio signal is modulated using periodic accent signals corresponding to the user's SRP. Therefore, the original accent rhythm of the speech contained in the captured audio signal is overridden by the user's SRP. Therefore, the voice is again adapted to the user's SRP to improve the user's voice perception.

[0022] Preferably, the artificial speech accent is selected to be similar to the natural speech accent. To this end, the amplitude of the processed sound signal is repeatedly increased within a predetermined time interval (TE), preferably within a time interval of 5 to 15 milliseconds, particularly 10 milliseconds. Specifically, each artificial accent involves an increase in the amplitude of the processed signal within said time interval.

[0023] In a suitable embodiment of the method, during OV analysis, the user's SRP is determined using a method known from EP 3 823 306 A1. Therefore, the same process is used to identify speech accents in the OV reference signal and speech accents in the speech of different speakers by determining at least one derivative of the amplitude and / or at least one derivative of the pitch of the OV reference signal and comparing this at least one derivative with a predefined standard.

[0024] However, in a similarly preferred embodiment of the invention, a different process is used to determine the user's SRP. Here, during the OV analysis process, the modulation depth of the (sound) amplitude modulation is determined, which is referred to as the "amplitude modulation depth" or "AMD" of the OV reference signal. Speech accents in the OV reference signal are determined by analyzing this AMD, which is determined within at least one predefined modulation frequency range. If the AMD meets predefined criteria, specifically exceeding a predefined threshold, then speech accents in the OV reference signal are identified (detected).

[0025] In two embodiments of the OV analysis process, the timing of speech accents identified in the OV reference signal and / or the time intervals between speech accents identified in the OV reference signal are determined. The user's SRP is derived from these determined times and / or time intervals.

[0026] In the foregoing, the term "(audio) amplitude modulation" refers to the change in the audio amplitude of the OV reference signal over time. The amplitude modulation depth (AMD) describes the normalized amplitude of the audio amplitude modulation:

[0027] Equation 1

[0028] Where A max and A min These are the maximum and minimum sound amplitudes of the OV reference signal within a given time window, respectively. If multiple modulation frequency ranges are analyzed, the AMD (i.e., a value of AMD) is calculated for each of the multiple modulation frequency ranges, and each of the multiple AMD values ​​is tested to see if it meets a predefined criterion, specifically whether it exceeds a corresponding threshold.

[0029] Since amplitude modulation is a time-dependent quantity, it can be described as a distribution of (modulation) frequencies. It is important to note that these modulation frequencies differ from the sound frequencies that form the OV reference signal. Preferably, the frequencies used to calculate A... max and A minThe time window is suitable for determining the modulation frequency range of the AMD. For example, the time window can be selected to closely correspond to the lower edge of the modulation frequency range. For example, if the lower edge of the modulation frequency range is 0.9 Hz (corresponding to a cycle time of 1.1 seconds), the assigned time window can be set to a value between 1.1 seconds and 1.5 seconds. Typically, the time window is selected such that it covers several (e.g., 1 to 5) oscillations of the sound amplitude within the corresponding frequency range.

[0030] In a preferred implementation of the aforementioned embodiments, the AMD of the OV reference signal is determined and analyzed for each of the three modulation frequency ranges, i.e.

[0031] The first modulation frequency range is -12 to 40 Hz, which corresponds to the typical rate of phonemes in speech.

[0032] The second modulation frequency range is -2.5-12 Hz, which corresponds to the typical rate of syllables in speech, and

[0033] The third modulation frequency range is -0.9 to 2.5 Hz, which corresponds to the typical rate of speech stress (i.e., emphasis) in speech.

[0034] Preferably, for each of these modulation frequency ranges, the corresponding AMD is compared with the corresponding threshold. Therefore, a speech accent is identified (detected) if the corresponding threshold is exceeded simultaneously in all three modulation frequency ranges (which may have the same or different values ​​for the three modulation frequency ranges). In other words, an event that exceeds the corresponding threshold only in one or two of the three modulation frequency ranges is not considered a speech accent. This allows for the identification of speech accents in the OV reference signal with very high sensitivity and selectivity.

[0035] In a simple yet effective embodiment of the invention, the time average (cycle time) between the identified speech accents of the OV reference signal is derived as a representation of the user's SRP. This represents the simplest case, where the SRP consists of a speech accent and the time interval to the next speech accent, such that in the user's speech stream, speech accents appear at approximately equal intervals (which is a fairly good approximation of real speech in most cases).

[0036] In more refined and precise embodiments, a more complex SRP is derived from the OV reference signal, which includes time or time intervals associated with multiple speech accents to more accurately represent the user's speech. Within the scope of this invention, statistical algorithms or methods for pattern recognition can be used to derive the user's SRP. For example, artificial intelligence such as artificial neural networks can be used.

[0037] In another embodiment of the method, speech accents in the OV reference signal are determined using only a finite frequency range of sound frequencies that form the OV reference signal. Preferably, this finite frequency range includes the low sound frequencies of the OV reference signal, such as the dominant 60 Hz to 5 kHz range in speech. Therefore, in the aforementioned embodiments, the method includes extracting the low sound frequency range from the OV reference signal and determining the user's SRP only from said low sound frequency range of the OV reference signal. Thus, potential interference from high-frequency noise contained in the OV reference signal to speech accent recognition can be avoided. Preferably, the low sound frequency range is extracted by dividing the OV reference signal into multiple sound bands and selectively using multiple low sound bands to identify speech accents in the OV reference signal.

[0038] In another embodiment of the method, the degree of difference between the user's SRP and the stresses contained in the captured audio signal is determined during the speech interval. Here, the enhancement process is performed only if the degree of difference exceeds a predefined threshold. In other words, the user's SRP is compared with the stresses contained in the captured audio signal. Here, the speech stresses in the captured audio signal are enhanced or induced only if the user's SRP is significantly different from the stresses contained in the captured audio signal. Otherwise, if the stresses contained in the captured audio signal are found to be very similar to or even identical to the user's SRP, the stresses contained in the captured audio signal are not touched; that is, no enhancement or induced speech stress is achieved. This embodiment is based on the experience that if the stresses contained in the captured audio signal are very similar to or even identical to the user's SRP, matching the captured audio signal to the user's SRP according to the present invention is neither necessary nor effective. Therefore, in the latter case, this matching should be omitted to avoid degrading sound quality. The degree of difference can be derived by determining the SRP of different speakers (i.e., speakers different from the user) based on the captured audio signal. Preferably, if applicable, this is done using the same method used to determine the user's SRP from the OV reference signal, and by associating the SRPs of different speakers with the user's SRP. If the SRP is represented by a single cycle time, the degree of difference can be determined by calculating the difference in cycle time between the user's SRP and the SRPs of different speakers. If the captured audio signal contains the speech of multiple different speakers, a separate SRP can be derived for each different speaker. Within the scope of this invention, the degree of difference can be defined in reverse, i.e., high values ​​for similar or identical SRPs and low values ​​for different SRPs. In this case, the enhancement process is only performed when the degree of difference is below a predefined threshold.

[0039] In a simple and suitable embodiment of the invention, the enhancement process is applied to all speech intervals, regardless of whether the captured sound signal contains the user's voice or the voice of a different speaker. However, depending on preference, the identified speech intervals are distinguished into self-voice intervals where the user is speaking and external voice intervals where at least one different speaker is speaking. In this case, preferably, the speech enhancement process is performed only during external voice intervals. In other words, during self-voice intervals, speech stress is not enhanced or caused. This embodiment reflects the experience that speech stress enhancement is unnecessary when the user is speaking because the user knows what he or she is saying, so perceiving his or her own voice is not a problem. By stopping the enhancement of speech stress during self-voice intervals, a processed sound signal containing a more natural voice of the user is provided to the user.

[0040] According to a second aspect of the invention, a hearing system having a hearing instrument (as previously described) is provided. The hearing instrument includes: an input converter configured to capture (raw) sound signals from the environment of the hearing instrument; a signal processor configured to process the captured sound signals to support the user's hearing (thereby providing a processed sound signal); and an output converter configured to transmit the processed sound signal to the user.

[0041] Typically, hearing systems are configured to automatically execute the method according to the first aspect of the invention. For this purpose, the system includes:

[0042] - A voice recognition unit configured to analyze captured sound signals to identify speech intervals;

[0043] - A speech enhancement unit configured to periodically change the amplitude of the processed sound signal according to a time pattern consistent with the user's SRP during the recognized speech interval.

[0044] For each embodiment or variation of the method according to the first aspect of the invention, there are corresponding embodiments or variations of the hearing system according to the second aspect of the invention. Therefore, the disclosures relating to the method are also applicable, in turn, to the hearing system.

[0045] Specifically, in a preferred embodiment, the hearing system further includes a derivation unit configured to determine at least one time derivative of the amplitude and / or pitch of the captured sound signal during the identified speech interval. Here, the speech enhancement unit is configured to temporarily increase the amplitude of the processed sound signal if at least one derivative meets a predefined criterion, particularly exceeding a predefined threshold, and if the time for meeting the predefined criterion is compatible with the user's SRP.

[0046] Additionally or alternatively, the speech enhancement unit is configured to superimpose artificial speech accents onto the captured audio signal by temporarily increasing the amplitude of the processed audio signal, such that the artificial speech accents match the user's SRP in a series of previous (artificial or natural) speech accents.

[0047] In another embodiment of the invention, the voice enhancement unit is configured to repeatedly increase the amplitude of the processed sound signal within a predetermined time interval, preferably within a time interval of 5 to 15 milliseconds, particularly within 10 milliseconds.

[0048] In another embodiment of the invention, the speech enhancement unit is configured to determine the degree of difference between the user's SRP and the repetitions contained in the captured audio signal during a speech interval; and to perform the enhancement process only if the degree of difference exceeds a predefined threshold.

[0049] In a further preferred embodiment, the hearing system includes a sound analysis unit configured to determine the user's SRP from an OV reference signal containing the user's speech. Preferably, the sound analysis unit is configured to...

[0050] - Determine speech accents from OV reference signals by analyzing the modulation depth of amplitude modulation of OV reference signals within at least one predefined modulation frequency range;

[0051] - Determine the timing of the identified speech accents in the OV reference signal and / or the time interval between the identified speech accents in the OV reference signal; and

[0052] - Derive the SRP from the determined time and / or time interval.

[0053] Preferably, the sound analysis unit is further configured as follows:

[0054] - Determine the modulation depth of the amplitude modulation of the OV reference signal for a first modulation frequency range of 12-40Hz, a second modulation frequency range of 2.5-12Hz, and a third modulation frequency range of 0.9-2.5Hz; and

[0055] - If the determined modulation depth exceeds a predefined threshold for each of the three modulation frequency ranges simultaneously, then speech accent is identified.

[0056] In another embodiment of the invention, the sound analysis unit is configured to derive the time average of the time intervals between speech accents as a representation of the SRP.

[0057] In another embodiment of the invention, the hearing system is configured to extract the low-frequency range of the OV reference signal. In this case, preferably, the sound analysis unit is configured to determine speech stress only from said low-frequency range of the OV reference signal.

[0058] Preferably, the signal processor is designed as a digital electronic device. It can be a single unit or composed of multiple subprocessors. The signal processor or at least one of the subprocessors can be a programmable device (e.g., a microcontroller). In this case, the functions mentioned above, or a portion thereof, can be implemented as software (especially firmware). Alternatively, the signal processor or at least one of the subprocessors can be a non-programmable device (e.g., an ASIC). In this case, the functions mentioned above, or a portion thereof, can be implemented as hardware circuitry.

[0059] In a preferred embodiment of the invention, the voice recognition unit, voice analysis unit, voice enhancement unit, and / or (if applicable) derivation unit are arranged within the hearing instrument. Specifically, each of these units may be designed as a hardware or software component of a signal processor, or as a separate electronic component. However, in other embodiments of the invention, at least one of the aforementioned units, particularly the voice recognition unit, voice analysis unit, and / or voice enhancement unit, or at least a functional portion thereof, may be implemented in an external electronic device such as a mobile phone.

[0060] In a preferred embodiment, the voice recognition unit includes a voice activity detection (VAD) module for detecting general voice activity (i.e., speech) and an OV detection (OVD) module for detecting the user's own voice. Attached Figure Description

[0061] Embodiments of the present invention will then be described with reference to the accompanying drawings, in which:

[0062] Figure 1 A schematic diagram of a hearing system including a hearing aid is shown. The hearing aid includes: an input converter arranged to capture sound signals from the environment of the hearing aid; a signal processor arranged to process the captured sound signals; and an output converter arranged to transmit the processed sound signals to a user.

[0063] Figure 2 It shows Figure 1 A schematic diagram of the functional structure of the signal processor in the hearing aid shown.

[0064] Figure 3 It shows the method for running Figure 1A flowchart of a method for a hearing aid, the method comprising: during speech enhancement, temporarily applying a gain to a captured sound signal and thus temporarily increasing the amplitude of the processed sound signal to enhance or induce speech accents in speech contained in the captured sound signal, wherein the speech accents are enhanced or induced in a time pattern consistent with a user-defined stress rhythm pattern (SRP).

[0065] Figure 4 The three synchronous plots over time show a series of speech accents identified in the foreign speech contained in the captured audio signal (top plot), binary time-related variables indicating time windows in which speech accents matching the user's SRP are expected to appear (middle plot), and the applied gain used to enhance the speech accents matching the user's SRP (bottom plot).

[0066] Figure 5 A flowchart is shown for the self-voice analysis process used to determine a user's SRP;

[0067] Figure 6 A flowchart illustrating an alternative embodiment of the self-voice analysis process for determining a user's SRP is shown; and

[0068] Figure 7 It shows including according to Figure 1 A schematic diagram of a hearing aid and a hearing system with a software application for controlling and programming the hearing aid, which is installed on a mobile phone.

[0069] Similar reference numerals denote similar parts, structures, and elements, unless otherwise specified. Detailed Implementation

[0070] Figure 1 A hearing system 2 is shown, which includes a hearing aid 4, i.e., a hearing instrument configured to support the hearing of a user with hearing loss, configured to be worn in or on one of the user's ears. Figure 1 As shown, as an example, hearing aid 4 can be designed as an behind-the-ear (BTE) hearing aid. Optionally, system 2 includes a second hearing aid (not shown) that is worn in or over the user's other ear to provide binaural support to the user.

[0071] The hearing aid 4 includes two microphones 6 as input converters and a receiver 8 as an output converter within the housing 5. The hearing aid 4 also includes a battery 10 and a signal processor 12. Preferably, the signal processor 12 includes programmable subunits (e.g., microprocessors) and non-programmable subunits (e.g., ASICs).

[0072] The signal processor 12 is powered by the battery 10, that is, the battery 10 provides the power supply voltage U to the signal processor 12.

[0073] During normal operation of hearing aid 4, microphone 6 captures airborne sound signals from the environment of hearing aid 2. Microphone 6 converts the airborne sound into an input audio signal I (also referred to as the "captured sound signal"), i.e., an electrical signal containing information about the captured sound. Input audio signal I is fed to signal processor 12. Signal processor 12 processes input audio signal I to provide directional sound information (beamforming), performs noise reduction and dynamic compression, and individually amplifies different spectral portions of input audio signal I based on the user's audiogram data to compensate for user-specific hearing loss. Signal processor 12 transmits output audio signal O (also referred to as the "processed sound signal") to receiver 8, i.e., an electrical signal containing information about the processed sound. Receiver 8 converts output audio signal O into processed airborne sound emitted into the user's ear canal by connecting receiver 8 to sound channel 14 of tip 16 of housing 5 and a flexible sound tube (not shown) connecting tip 16 to earpiece (which is inserted into the user's ear canal).

[0074] like Figure 2 As shown, the signal processor 12 includes a voice recognition unit 18, which includes a voice activity detection (VAD) module 20 and a self-voice detection (OVD) module 22. Preferably, modules 20 and 22 are both designed as software components installed in the signal processor 12.

[0075] VAD module 20 typically detects the presence of sound (i.e., speech, independent of a specific speaker) in the input audio signal I, while OVD module 22 specifically detects the presence of the user's own voice in the input audio signal I. Preferably, modules 20 and 22 apply VAD and OVD techniques known in the art, such as those from US2013 / 0148829A1 or WO 2016 / 078786A1. By analyzing the input audio signal I (and thus the captured sound signal), VAD module 20 and OVD module 22 identify speech intervals, where the input audio signal I contains speech, which are distinguished (subdivided) into self-voice intervals (OV intervals) where the user speaks and foreign-voice intervals (FV intervals) where at least one different speaker speaks when the user is silent.

[0076] In addition, the hearing system 2 includes a derivation unit 24, a speech enhancement unit 26, and a sound analysis unit 28.

[0077] The derivation unit 24 is configured to derive the pitch P (i.e., fundamental frequency) from the input audio signal I as a time-dependent variable. The derivation unit 24 is also configured to apply a moving average to the measured value of the pitch P, for example, by applying a time constant of 15 milliseconds (i.e., the size of the time window used for averaging), to derive the first (time) derivative D1 and the second (time) derivative D2 of the time average of the pitch P.

[0078] For example, in a simple and efficient implementation, a periodic time series of the time average of pitch P is given by ..., AP[n-2], AP[n-1], AP[n], ..., where AP[n] is the current value, and AP[n-2] and AP[n-1] are previously determined values. Then, the current value D1[n] and the previous value D1[n-1] of the first derivative D1 can be determined as...

[0079] D1[n] = AP[n] – AP[n-1] = D1,

[0080] D1[n-1] = AP[n-1] – AP[n-2],

[0081] The current value of the second derivative D2, D2[n], can be determined as follows:

[0082] D2[n] = D1[n] – D1[n-1] = D2.

[0083] The speech enhancement unit 26 is configured to analyze derivatives D1 and D2 relative to a criterion described in more detail later in order to identify speech accents in the input audio signal I (and thus the captured sound signal). Furthermore, the speech enhancement unit 26 is configured to temporarily apply an additional gain G to the input audio signal I (in its initial or partially processed form), and thus increase the amplitude of the processed sound signal O if derivatives D1 and D2 satisfy the criterion (indicating speech accents).

[0084] Preferably, the derivation unit 24, the speech enhancement unit 26, and the sound analysis unit 28 are designed as software components installed in the signal processor 12.

[0085] During normal operation of the hearing aid 4, the sound recognition unit 18, namely the VAD module 20 and the OVD module 22, the derivation unit 24, and the speech enhancement unit 26 interact to perform... Figure 3 Method 30 is shown.

[0086] In the (speech recognition) step 31 of the method, the voice recognition unit 18 analyzes the input audio signal I at FV intervals, that is, it checks whether the VAD module 20 returns a positive result (which means that speech is detected in the input audio signal I), while the OVD module 22 returns a negative result (which means that there is no user's own voice in the input audio signal I).

[0087] If an FV interval (Y) is detected, the sound recognition unit 18 triggers the derivation unit 24 to execute the next step 32. Otherwise (N), step 31 is repeated.

[0088] In step 32, the derivation unit 24 derives the pitch P of the captured sound from the input audio signal I and applies a time average to the pitch P as described above. In the subsequent (derivation) step 34, the derivation unit 24 derives the first derivative D1 and the second derivative D2 of the time average of the pitch P.

[0089] Subsequently, the derivation unit 24 triggers the speech enhancement unit 26 to execute the speech enhancement process 36. Figure 3 In the example shown, the process is broken down into four steps: 38, 40, 42, and 44.

[0090] In step 38, the speech enhancement unit 26 analyzes the derivatives D1 and D2 as described above to identify speech stress. If a speech stress is identified (Y), the speech enhancement unit 26 proceeds to step 40. Otherwise (N), that is, if no speech stress is identified, the speech enhancement unit 26 triggers the voice recognition unit 18 to execute step 31 again.

[0091] Preferably, the speech enhancement unit 26 uses one of the algorithms described in EP 3 823 306 A1 to identify speech stresses of different speakers in the input audio signal I, wherein the aforementioned criteria for identifying speech stresses involve comparing a first derivative D1 of the time-averaged pitch P with a threshold, which is further influenced by a second derivative D2.

[0092] According to the first algorithm, the speech enhancement unit 26 checks whether the first derivative D1 exceeds a threshold. If yes (Y), the speech enhancement unit 26 proceeds to step 40. Otherwise (N), the speech enhancement unit 26 triggers the voice recognition unit 18 to execute step 31 again. The threshold is offset (changes) according to the second derivative D2, as described in EP 3 823 306 A1.

[0093] According to the second algorithm, the speech enhancement unit 26 multiplies the first derivative D1 by a variable weight factor determined according to the second derivative D2, as described in EP 3 823 306 A1. Subsequently, the speech enhancement unit 26 checks whether the weighted first derivative D1 exceeds a threshold. If yes (Y), the speech enhancement unit 26 proceeds to step 40. Otherwise (N), the speech enhancement unit 26 triggers the voice recognition unit 18 to execute step 31 again.

[0094] If step 38 produces a positive result (Y), then the speech enhancement unit 26 executes step 40. In step 40, the speech enhancement unit 26 checks whether the current time (i.e., the time point at which the speech accent was identified in the previous step 38) matches the user's predefined stress rhythm pattern (SRP). For this purpose, for example, the speech enhancement unit 26 may check whether the binary time-related variable V has a value of 1 (V=1?). If yes (Y), it indicates that the identified speech accent matches the user's personal stress rhythm, and the speech enhancement unit 26 proceeds to step 42. Otherwise (N), the speech enhancement unit 26 triggers the voice recognition unit 18 to execute step 31 again via step 44, which is described later.

[0095] In step 42, the speech enhancement unit 26 temporarily applies an additional gain G to the captured audio signal. Therefore, for a predefined time interval (referred to as the enhancement interval TE), the amplitude of the processed audio signal O is increased, thereby enhancing the recognized speech accents. After the enhancement interval TE expires, the additional gain G decreases to 1 (0 dB). The speech enhancement unit 26 triggers the voice recognition unit 18 to execute step 31 via step 44, thereby restarting the process based on… Figure 3 The method. As previously described, an additional gain G can be applied to the captured audio signal at any stage of signal processing. Therefore, it can be applied to the input audio I initially captured by microphone 6, but it can also be applied to audio signals captured after one or more preceding signal processing steps.

[0096] As disclosed in EP 3 823 306 A1, the additional gain G can be, for example,

[0097] - Gradually increasing and decreasing (i.e., as a bivariate function of time) or

[0098] - With a linear or non-linear dependence over time, it gradually increases and continuously decreases, or

[0099] - It increases and decreases continuously with a linear or non-linear dependence over time.

[0100] Initially, variable V is preset to a constant of 1 (V=1). Therefore, when step 40 is executed for the first time within the FV interval, a positive result (Y) is always produced, and speech enhancement unit 26 always proceeds to step 42.

[0101] Subsequently, in step 44, the speech enhancement unit 26 modifies the variable V to indicate the time window for the expected future speech stress (based on the user's SRP). Within each time window, the variable V is assigned a value of 1 (V=1). Outside the time window, the variable V is assigned a value of 0 (V=0). Figure 3 In the example shown, the user's SRP is represented by the average time interval between consecutive speech stresses of the user's own voice (hereafter referred to as the cycle time C). Therefore, time windows are selected to match the cycle time C plus or minus its confidence interval ΔC: from the point in time at which step 42 is performed, the first time window will begin at C - ΔC and end at C + ΔC. Similarly, the second time window will begin at 2·C - ΔC and end at 2·C + ΔC, the third time window will begin at 3·C - ΔC and end at 3·C + ΔC, and so on.

[0102] If step 31 produces a negative result (N), indicating the end of the FV interval or its non-existence, then variable V is reset to the constant 1 (V=1).

[0103] Variable V pairs Figure 3 The effect of method 30 shown is in Figure 4 As shown in the image.

[0104] As an example, Figure 4 The figure above illustrates a series of events occurring over time t, representing the input audio signal I and thus the captured sound signal. At time t1, the FV interval begins. At times t2, t3, t4, t5, and t6, five consecutive speech accents of the foreign sound are identified in the input audio signal I. At time t7, the FV interval ends.

[0105] exist Figure 4 The middle figure shows the time dependency of variable V. Figure 4 The figure below shows the corresponding time dependence of the additional gain G.

[0106] It can be seen that before the first speech stress is detected at time t2, the variable V is pre-set to a constant 1 (V=1). Therefore, at time t2, the first execution based on... Figure 4 Steps 40 and 42 of the method.

[0107] In step 42, at time t2, the gain G is temporarily increased to enhance the first speech accent.

[0108] In step 44, the variable V is modified to indicate a series of time windows as described above (shown as shaded areas). The first time window begins at time t2 + C - ΔC, and the second time window begins at time t2 + 2·C - ΔC. The duration of each time window is 2·ΔC. As shown, within each time window, the value of the variable is 1 (V=1), while outside the time window, the variable V is set to a constant of zero (V=0).

[0109] from Figure 4 It can be seen that the first time window passed without recognizing any further speech stress. In fact, the second speech stress was recognized at time t3 between the first and second time windows. Since the variable V was set to zero (V=0) at time t3, step 40 produces a negative result (N). Therefore, step 42 is not executed, and the second speech stress is not amplified.

[0110] The third speech accent is detected at time t4 within the second time window. Therefore, step 42 is executed. Subsequently, in step 44, variable V is reset to zero and modified to indicate the appropriate time window (where the first time window begins at time t4 + C - ΔC). In step 42, at time t4, the gain G is temporarily increased to enhance the third speech accent.

[0111] For the fourth and fifth speech stresses, the process is repeated at times t5 and t6.

[0112] At time t7, at the end of the foreign speech interval, step 31 produces a negative result. Therefore, variable V is reset to the constant 1 (V=1).

[0113] Optionally, in Figure 3 In the improved version of method 30 shown, the speech enhancement unit 26 ( Figure 4 Artificial speech accents are created at the end of any time window in which no (natural) speech accent is recognized in the input audio signal I. Figure 4 In the example, an artificial speech accent 46 can be created at time t2 + C + ΔC at the end of the first time window to fill the gap between the natural speech accents recognized at times t2 and t4 in a time pattern consistent with the user's SRP. Figure 4 As shown by the dashed line, artificial speech accent 46 is created by temporarily increasing the gain G in the same manner as at times t2, t4, t5, and t6.

[0114] Preferably, the variables characterizing the user's SRP, namely the cycle time C and the confidence interval ΔC, are determined by the hearing system 2 during the setup process prior to the normal operation of the hearing aid 4. For this purpose, the sound recognition unit 18, the derivation unit 24, and the sound analysis unit 28 interact to perform... Figure 5The self-voice analysis process (OV analysis process) shown in Figure 50.

[0115] In the OV recognition step 51 of the OV analysis process 50, the voice recognition unit 18 analyzes the input audio signal I at OV intervals, that is, checks whether the OVD module 22 returns a positive result (indicating that the user's own voice is detected in the input audio signal I). If yes (Y), the voice recognition unit 18 triggers the derivation unit 24 to execute step 52. Otherwise (N), step 51 is repeated.

[0116] In step 52 and subsequent steps 54 and 56, respectively, similar to Figure 3 The corresponding steps 32, 34, and 38 of the method involve determining the pitch P of the user's own voice (step 52), deriving the first and second derivatives D1 and D2 of the time-averaged pitch P by the derivation unit 24 (derivation step 54), and identifying speech stresses in the user's own voice by the sound analysis unit 28 (step 56). If the sound analysis unit 28 (Y) identifies a speech stress, the speech enhancement unit 26 proceeds to step 58. Otherwise (N), i.e., if no speech stress is identified, the sound analysis unit 28 triggers the sound recognition unit 18 to execute step 51 again.

[0117] In step 58, the sound analysis unit 28 determines and statistically evaluates the time required to identify speech stress in the user's self-voice, and determines the cycle time C and confidence interval ΔC as representations of the user's SRP. For example, the confidence interval ΔC can be determined as the standard deviation or statistical range of the measured value of the cycle time C of the user's self-voice. The cycle time C and confidence interval ΔC are stored in the memory of the hearing aid 4 for later use during normal operation of the hearing aid 4.

[0118] The process 50 terminates when a sufficient number of speech stresses (e.g., 1000) of the user's self-voice have been identified and evaluated. Therefore, the sound signal captured by the hearing instrument 4 during the OV interval is used as the OV reference signal to derive the user's SRP.

[0119] exist Figure 5 In a variation of process 50 shown, steps 52 to 58 are performed during OV and FV intervals, particularly during normal operation of the hearing aid 4. Therefore, in step 58, the sound analysis unit 28 is applied to both the external sound and the user's own voice, and separate values ​​for the cycle time C of the external sound and the user's own voice are derived. Thus, the sound analysis unit 28 determines the user's SRP and the external speaker's SPR. Figure 3In a corresponding variation of method 30, for example, in step 44 or 31, the speech enhancement unit 26 compares the user's SRP and the outside speaker's SPR by comparing the difference between the cycle time C values ​​of the outside voice and the self-voice with a threshold, respectively. In this case, if it is found that the user's SRP and the outside speaker's SPR are sufficiently different, i.e., if the difference between the cycle time C values ​​of the outside voice and the self-voice exceeds the threshold, then only the speech stress of the outside voice is enhanced (step 42). Otherwise, the variable V is set to 0, which causes step 42 to be skipped.

[0120] Optionally, the sounds contained in the input audio signal I during the FV interval are analyzed to distinguish the voices of multiple different speakers (if any). In this case, the SRP (i.e., the value of the cycle time C) is determined separately for each individual different speaker.

[0121] Figure 6 An alternative OV analysis procedure 60 is shown for deriving the variables characterizing the user's SRP, namely the cycle time C and the confidence interval ΔC.

[0122] In the OV recognition step 61 of this process, similar to Figure 5 In step 51 of the process, the voice recognition unit 18 analyzes the input audio signal I at OV intervals, that is, it checks whether the OVD module 22 returns a positive result (indicating that the user's own voice is detected in the input audio signal I). If yes (Y), the voice recognition unit 18 triggers the voice analysis unit 28 to execute step 62. Otherwise (N), step 61 is repeated.

[0123] In step 62, the sound analysis unit 28 determines the amplitude modulation A of the input audio signal I (i.e., the time-dependent envelope of the input audio signal I). Furthermore, the sound analysis unit 64 divides the amplitude modulation A into three modulation bands (bands of modulation frequencies), namely...

[0124] - The first modulation band includes modulation frequencies in the range of 12-40 Hz, corresponding to the typical rates of phonemes in speech.

[0125] - The second modulation band includes modulation frequencies in the range of 2.5-12 Hz, corresponding to the typical rate of syllables in speech, and

[0126] - The third modulation band includes modulation frequencies in the range of 0.9–2.5 Hz, corresponding to the typical rate of speech stress (i.e., emphasis) in speech.

[0127] For each of the three modulation bands, in step 64, the sound analysis unit 28 determines the respective modulation depths M1, M2, and M3 by evaluating the maximum and minimum sound amplitudes within a time window and for each modulation band, according to Equation 1. For example, the time window is set to 84 milliseconds for the first modulation band, 400 milliseconds for the second modulation band, and 1100 milliseconds for the third modulation band.

[0128] In step 66, the modulation depths M1, M2, and M3 are compared with corresponding thresholds to identify speech accents. If the modulation depths M1, M2, and M3 of all modulation bands simultaneously exceed the corresponding threshold (Y), the sound analysis unit 28 identifies a speech accent and proceeds to step 68. Otherwise (N), i.e., if no speech accent is identified, the sound analysis unit 28 triggers the sound recognition unit 18 to execute step 61 again.

[0129] In step 68, similar to Figure 5 In step 58, the sound analysis unit 28 determines and statistically evaluates the time required to identify speech stress in the user's own voice. Specifically, the cycle time C and confidence interval ΔC are determined as representations of the user's SRP, as previously described. The cycle time C and confidence interval ΔC are stored in the memory of the hearing aid 4 for later use during normal operation of the hearing aid 4.

[0130] The process ends 60 when enough voice accents of the user's own speech have been identified and evaluated (e.g., 1000 voice accents).

[0131] According to respectively Figure 5 or Figure 6In a more improved and precise embodiment of processes 50 and 60, a more complex SRP for the user is derived from the audio input signal I during the OV interval, which includes multiple time intervals between consecutive speech accents. For this purpose, for example, the time intervals between speech accents identified in one of steps 56 or 66 are divided into N groups of consecutive speech accents, where N is an integer whose value varies (N=2, 3, 4…) to find the best matching pattern. For example, the time intervals of identified speech accents are split into groups of two consecutive speech accents, groups of three consecutive speech accents, groups of four consecutive speech accents, etc. For each N value, these groups are compared to each other. The group most similar to each other is selected to derive the SRP, for example, by averaging over the corresponding time or time interval of the selected group. For example, if the analysis reveals that the group of three consecutive speech accents (N=3) is more similar to each other than the group of two consecutive speech accents (N=2) and the group of four consecutive speech accents (N=4), then the group of three consecutive speech accents is selected to derive the SRP. In this case, the user's SRP can be derived by averaging over the corresponding first time interval in the selected group, by averaging over the corresponding second time interval in the selected group, and by averaging over the corresponding third time interval in the selected group. In this case, the SRP is represented by the average first time interval between the first and second speech accents of the SRP, the average second time interval between the second and third speech accents of the SRP, and the average third time interval after the third speech accent of the SRP.

[0132] Within the scope of this invention, other static algorithms or methods of pattern recognition can be used to derive a user's SRP. For example, artificial intelligence such as artificial neural networks can be used.

[0133] In another embodiment, the input audio signal I (particularly the input audio signal I captured during the OV interval) is divided into multiple (audio) frequency bands before being fed to the sound analysis unit 28. In this case, preferably, the low-frequency range of the input audio signal I (including the lower subset of the audio bands) is selectively analyzed in the OV analysis process 50 or 60. In other words, one or more high-frequency bands are excluded from the OV analysis process 50 or 60 (i.e., not analyzed).

[0134] Figure 7 Another embodiment of the hearing system 2 is shown, wherein the hearing system 2 includes a hearing aid 4 as described above and a software application (hereinafter referred to as "hearing application" 70) installed on a user's mobile phone 72. Here, the mobile phone 72 is not part of the system 2. Rather, it is used by the hearing system 2 only as an external resource to provide computing power and memory.

[0135] Hearing aid 4 and hearing application 70 exchange data via wireless link 74, for example, based on the Bluetooth standard. To this end, hearing application 70 accesses the wireless transceiver of mobile phone 72, particularly the Bluetooth transceiver (not shown), to send data to and receive data from hearing aid 4.

[0136] According to Figure 7 In some embodiments, some elements or functions of the aforementioned hearing system 2 (instead of the hearing aid 2) are implemented in the hearing application 70. For example, at least one functional portion of the speech enhancement unit 26 configured to perform step 38 is implemented in the hearing application 70. Additionally or alternatively, the sound analysis unit 28 may be implemented in the hearing application 72.

[0137] Those skilled in the art will understand that various changes and / or modifications can be made to the invention as illustrated in the specific examples without departing from the spirit and scope of the invention as broadly described herein. Therefore, the present examples should be considered illustrative in all respects, not restrictive.

[0138] List of reference numerals

[0139] 2 (Hearing) System

[0140] 4 hearing aids

[0141] 5. Shell

[0142] 6 microphones

[0143] 8 receivers

[0144] 10 batteries

[0145] 12 signal processors

[0146] 14 audio channels

[0147] 16 tips

[0148] 18 voice recognition units

[0149] 20. Sound Activity Detection Module (VAD Module)

[0150] 22. Self-Voice Detection Module (OVD Module)

[0151] 24 Derivation Units

[0152] 26 speech enhancement units

[0153] 28 sound analysis units

[0154] 30 methods

[0155] 31 (Speech Recognition) Steps

[0156] 32 steps

[0157] 34 (Derivation) Steps

[0158] 36 (Speech Enhancement) Process

[0159] 38 steps

[0160] 40 steps

[0161] 42 steps

[0162] 44 steps

[0163] 46 (Artificial) Speech Stress

[0164] 50. Self-Voice Analysis Process (OV Analysis Process)

[0165] 51 (OV recognition) steps

[0166] 52 steps

[0167] 54 (Derivation) Steps

[0168] 56 steps

[0169] 58 steps

[0170] 60. Self-Voice Analysis Process (OV Analysis Process)

[0171] 61 (OV recognition) steps

[0172] 62 steps

[0173] 64 steps

[0174] 66 steps

[0175] 68 steps

[0176] 70 Hearing App

[0177] 72 mobile phone

[0178] 74 wireless links

[0179] ΔC confidence interval

[0180] t time

[0181] t1 time

[0182] t2 time

[0183] t3 time

[0184] t4 time

[0185] t5 time

[0186] t6 time

[0187] t7 time

[0188] Amplitude modulation

[0189] C loop time

[0190] I input audio signal

[0191] D1 (first derivative)

[0192] D2 (second) derivative

[0193] G gain

[0194] M1 modulation depth

[0195] M2 modulation depth

[0196] M3 modulation depth

[0197] O outputs audio signal

[0198] P pitch

[0199] TE Enhanced Interval

[0200] U (electric) power supply voltage

[0201] V variable.

Claims

1. A method for operating a hearing instrument (4), the method comprising: - Capture sound signals from the environment of the hearing instrument (4); - Process the captured audio signals; - Output the processed sound signal to the user of the hearing instrument (4); The method further includes: - In the speech recognition step (31), the captured sound signal is analyzed to identify speech intervals, wherein the captured sound signal contains speech; and - In the speech enhancement process (36) performed during the identified speech intervals, the amplitude of the processed sound signal is periodically changed. Its features are, In the speech enhancement process (36), the amplitude of the processed sound signal is periodically changed according to a time pattern consistent with the user's stress rhythm pattern, which is a personal stressed rhythm of the user's speech used to construct and emphasize the speech and contains speech stress, i.e., a single peak in amplitude and / or pitch, the speech stress having a duration of 5 to 15 milliseconds and occurring at time intervals of more than 400 milliseconds between each other. The method further includes determining the user's stress rhythm pattern based on a self-voice reference signal containing the user's speech in a self-voice analysis process (50, 60), and the speech enhancement process (36) includes superimposing an artificial speech stress (46) on the captured sound signal by temporarily increasing the amplitude of the processed sound signal, such that the artificial speech stress (46) matches the user's stress rhythm pattern in a series of previous speech stresses.

2. The method according to claim 1, It also includes, in the derivation step (34) performed during the identified speech interval, determining at least one time derivative (D1, D2) of the amplitude and / or pitch (P) of the captured sound signal, wherein, If at least one derivative (D1, D2) satisfies a predefined criterion, and if the time for satisfying the predefined criterion is compatible with the user’s repetition rhythm pattern, then the speech enhancement process (36) includes temporarily increasing the amplitude of the processed sound signal.

3. The method according to claim 1 or 2, The speech enhancement process (36) includes repeatedly increasing the amplitude of the processed sound signal within a predetermined time interval (TE).

4. The method according to claim 1 or 2, in, The self-voice analysis process (60) includes: - Determine the modulation depth (M1, M2, M3) of the amplitude modulation (A) of the self-sound reference signal within a predefined modulation frequency range. - Speech accents in a self-speech reference signal are identified by analyzing the modulation depths (M1, M2, M3), wherein a speech accent is identified if the modulation depths (M1, M2, M3) meet predefined criteria. - Determine the timing of speech stresses identified in the self-voice reference signal and / or the time interval between identified speech stresses in the self-voice reference signal; and - Derive the user's repetition rhythm pattern from the determined time and / or time interval.

5. The method according to claim 4, The steps for identifying speech stress in a self-voice reference signal include: - For each of the first modulation frequency range of 12-40 Hz, the second modulation frequency range of 2.5-12 Hz, and the third modulation frequency range of 0.9-2.5 Hz, determine the modulation depth (M1, M2, M3) of the amplitude modulation (A) of the self-sound reference signal; and - If the determined modulation depth (M1, M2, M3) exceeds the corresponding predefined threshold of each of the three modulation frequency ranges mentioned above, then the speech accent of the self-voice reference signal is identified.

6. The method according to claim 1, The average time interval between speech stresses derived from the user's own voice reference signal is used as a representation of the user's stress rhythm pattern.

7. The method according to claim 1 or 2, in, The self-voice analysis process also includes: - Extract the low-frequency range of the self-voice reference signal; and - The user's repetition rhythm pattern is determined only from the low sound frequency range.

8. The method according to claim 1 or 2, - Wherein, during the speech interval, the degree of difference between the stress rhythm pattern of the self-sound reference signal and the stress contained in the captured sound signal is determined; and - Wherein, the speech enhancement process is performed only when the degree of difference exceeds a predefined threshold (36).

9. A hearing system (2) having a hearing instrument (4), said hearing instrument (4) comprising: - An input converter (6) is arranged to capture sound signals from the environment of the hearing instrument (4); - Signal processor (12), which is arranged to process the captured sound signal; and - Output converter (8), which is arranged to transmit processed sound signals to the user of the hearing instrument. The hearing system (2) also includes: - A voice recognition unit (18) configured to analyze a captured sound signal to identify speech intervals, wherein the captured sound signal contains speech; and - A speech enhancement unit (26) configured to periodically change the amplitude of the processed sound signal during a speech enhancement process (36) performed during the recognized speech interval. Its features are, The speech enhancement unit (26) is configured such that, during the speech enhancement process (36), the amplitude of the processed sound signal is periodically changed according to a time pattern consistent with the user's stress rhythm pattern, the stress rhythm pattern being a personal stressed rhythm of the user's speech used to construct and emphasize the speech and containing speech stress, i.e., a single peak in amplitude and / or pitch, the speech stress having a duration of 5 to 15 milliseconds and occurring at time intervals of more than 400 milliseconds between each other, wherein the hearing system (2) further includes a sound analysis unit (28) configured to determine the user's stress rhythm pattern based on a self-sound reference signal containing the user's speech, and wherein the speech enhancement unit (26) is configured to superimpose artificial speech stress on the captured sound signal by temporarily increasing the amplitude of the processed sound signal, such that the artificial speech stress matches the user's stress rhythm pattern in a series of previous speech stresses.

10. The hearing system (2) according to claim 9. It also includes a derivation unit (24) configured to determine at least one time derivative (D1, D2) of the amplitude and / or pitch (P) of the captured sound signal during the identified speech interval, wherein, If at least one derivative (D1, D2) satisfies a predefined criterion, and if the time for satisfying the predefined criterion is compatible with the user's repetition rhythm pattern, then the speech enhancement unit (26) is configured to temporarily increase the amplitude of the processed sound signal.

11. The hearing system (2) according to claim 9 or 10. in, The speech enhancement unit (26) is configured to repeatedly increase the amplitude of the processed sound signal within a predetermined time interval (TE).

12. The hearing system (2) according to claim 9 or 10. in, The sound analysis unit (28) is configured as follows: Determine the modulation depth (M1, M2, M3) of the amplitude modulation (A) of the self-sound reference signal within at least one predefined modulation frequency range. - Speech accents are identified from the self-voice reference signal by analyzing the modulation depths (M1, M2, M3), wherein speech accents are identified if the modulation depths (M1, M2, M3) meet predefined criteria. - Determine the time of the recognized speech accent in the self-voice reference signal and / or the time interval between the recognized speech accents in the self-voice reference signal; and - Derive the user's repetition rhythm pattern from the determined time and / or time interval.

13. The hearing system (2) according to claim 12. in, The sound analysis unit (28) is configured as follows: - For a first modulation frequency range of 12-40Hz, a second modulation frequency range of 2.5-12Hz, and a third modulation frequency range of 0.9-2.5Hz, determine the modulation depth (M1, M2, M3) of the amplitude modulation (A) of the self-sound reference signal; and - If the determined modulation depth exceeds the corresponding predefined threshold of each of the three modulation frequency ranges mentioned above, then the speech accent of the self-voice reference signal is identified.

14. The hearing system (2) according to claim 9 or 10. in, The sound analysis unit (28) is configured to derive the time average of the time interval between speech stresses in its own sound reference signal as a representation of the stress rhythm pattern.

15. The hearing system (2) according to claim 9 or 10. The hearing system is configured to extract the low frequency range of the self-voice reference signal, wherein the sound analysis unit (28) is configured to determine the speech stress of the self-voice reference signal only from the low frequency range of the self-voice reference signal.

16. The hearing system (2) according to claim 9 or 10. in, The speech enhancement unit (26) is configured to - During speech intervals, determine the degree of difference between the user's stress rhythm pattern and the stress contained in the captured audio signal; and - The speech enhancement process is performed only when the degree of difference exceeds a predefined threshold (36).

Citation Information

Patent Citations

  • Hearing aid having an improved speech intelligibility by means of frequency selective signal processing, and a method for operating such a hearing aid

    EP1101390B1

  • Hearing apparatus with speaker activity detection and method for operating a hearing apparatus

    US20130148829A1

  • Method and apparatus for fast recognition of a user's own voice

    WO2016078786A1

  • A hearing system comprising a hearing instrument and a method for operating the hearing instrument

    EP3823306A1

  • Hearing aid systems and methods

    WO2021136962A1