Method for operating a hearing aid

DE102024205064B4Active Publication Date: 2025-09-11SIVANTOS PTE LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102024205064
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-05-31
Publication Date
2025-09-11
Estimated Expiration
2044-05-31

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method for operating a hearing aid (2), comprising - a receiving unit (4) for receiving speech information and converting it into a speech signal (S), - a signal processing unit (12) for processing the speech signal (S) and generating an output signal (A), and - an output unit (18) for outputting an output signal (A) to a hearing aid user, - wherein speech information is converted into a speech signal (S), - wherein the speech signal (S) is processed by the signal processing unit (12) into an output signal (A) for output to the hearing aid user, - wherein at least one emotion-correlated feature (F1, ..., F 18 , F1', ..., F 18 ') is determined, - where the emotion-correlated features (F1, ..., F 18, F1', ..., F 18 ') are compared with each other and a comparison result is determined, and - wherein the signal processing unit (12) is adjusted depending on the comparison result.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for operating a hearing aid. The invention further relates to a hearing aid for implementing the method.

[0002] A hearing aid is generally defined as an electronic device that supports the hearing of a person wearing the hearing aid. In particular, the invention relates to a hearing aid that is designed to fully or partially compensate for the hearing loss of a hearing-impaired user. Such a hearing aid is also referred to as a "hearing aid" (HA). There are also hearing aids that protect or improve the hearing of users with normal hearing, for example, to enable improved speech understanding in complex listening situations. Such devices are also referred to as "personal sound amplification products" (PSAPs for short). Finally, the term "hearing aid" as used here also includes headphones worn on or in the ear (wired or wireless, with or without active noise cancellation), headsets, etc., as well as implantable hearing aids, such as cochlear implants.A hearing aid can also be part of an AR (Augmented Reality) or VR (Virtual Reality) system to deliver the acoustic information of a virtual sound source to the user.

[0003] Hearing aids in general, and hearing aid devices in particular, are usually designed to be worn on the user's head, particularly in or on one ear, particularly as behind-the-ear (BTE) or in-the-ear (ITE) devices. Due to their internal structure, hearing aids typically have at least one output transducer that converts an input audio signal into a signal perceivable by the user as sound, and then outputs the latter to the user.

[0004] In most cases, the output transducer is designed as an electro-acoustic transducer that converts the (electrical) output audio signal into airborne sound, which is then emitted into the user's ear canal. In a hearing aid worn behind the ear, the output transducer, also known as the "receiver," is usually integrated outside the ear in a hearing aid housing. In this case, the sound emitted by the output transducer is guided into the user's ear canal via a sound tube. Alternatively, the output transducer can be located in the ear canal, and thus outside the housing worn behind the ear. Such hearing aids are also referred to as "Receiver In Channel" devices.Hearing aids worn in the ear, which are so small that they do not protrude beyond the ear canal, are also called CIC devices (after the English term “Completely in Canal”).

[0005] In other designs, the output transducer can also be an electromechanical transducer that converts the output audio signal into structure-borne sound (vibrations), which is then transmitted, for example, to the user's skull. There are also implantable hearing aids, particularly cochlear implants, and hearing aids whose output transducers directly stimulate the user's auditory nerve.

[0006] In addition to the output transducer, a hearing aid often has at least one (acousto-electrical) input transducer. During operation of the hearing aid, the or each input transducer picks up airborne sound from the hearing aid's environment and converts this airborne sound into an input audio signal (i.e., an electrical signal that conveys information about the ambient sound). This input audio signal - also referred to as the "picked up sound signal" - is regularly output to the user in its original or processed form, e.g., to implement a so-called transparency mode in headphones, for active noise cancellation, or - e.g., in the case of a hearing aid - to achieve improved noise perception for the user.

[0007] In addition, a hearing aid often has a signal processing unit (signal processor). The signal processing unit processes the or each input audio signal (i.e., modifies its sound information). The signal processing unit then outputs a correspondingly processed audio signal (also referred to as the "output audio signal" or "modified sound signal") to the output transducer and / or to an external device.

[0008] Hearing aids offer various additional (hearing / hearing aid) functions, for example in the area of ​​signal processing, which can improve the hearing benefit for a hearing aid wearer (HAW). Examples of such functions could be: Own Voice Detection (OVD), speech recognition or Voice Activity Detection (VAD), Active Noise Reduction (ANR), Active Occlusion Reduction (AOR); streaming of audio information (e.g., music); detection of various body signals (e.g., fitness); detection of specific events and reactions to them (e.g., fall detection - if a user falls, an alarm is sent), etc. Such functions are often implemented as software or processing algorithms of the signal processing unit.

[0009] When using hearing systems or hearing aids, handling conversational situations is one of the core problems. This is primarily due to the fact that hearing system users often receive important information during face-to-face conversations. Another important piece of information is the other person's emotions.

[0010] As is known, for example, from BW Schuller, "Speech emotion recognition: two decades in a nutshell, benchmarks and ongoing trends," Communications of the ACM, Volume 61, Issue 5, pp. 90-99, a speaker's emotions can be detected and classified in speech signals. For example, there are several relevant acoustic features that can be extracted from the speech signal. Furthermore, textual content or emotional keywords can also be used to support acoustic (emotion) analysis.

[0011] To understand the patterns of vocal expression of various emotions and other affective dispositions and processes, it is possible to extract acoustic parameters from the speech signal. The underlying theoretical assumption is that affective processes differentially alter the autonomic excitation and tension of the striatal muscles, thereby influencing the voice and speech production at the phonatory and articulatory levels. These changes can be assessed using various parameters of the acoustic waveform.

[0012] During operation, a hearing aid utilizes a variety of signal programming algorithms to compensate for the hearing loss, but also to provide additional support algorithms for optimal speech understanding and / or comfortable listening. All of these manipulations of the original input signal can influence the acoustic characteristics important for emotion recognition. Compensating for hearing loss with a hearing aid can therefore be accompanied by a deterioration in emotion recognition. As a result, the hearing aid wearer may not have access to all the signal information needed for emotion recognition in a conversational situation.

[0013] In the publication Lee et al., “Deep Representation Learning for Affective Speech Signal Analysis and Processing: Preventing unwanted signal disparities”, IEEE Signal Processing Magazine, vol. 38, no. 6, pp. 22-38, Nov. 2021 (doi: 10.1109 / MSP.2021.3105939), a method is described in which a speech signal is processed into an output signal in such a way that an emotion in the output signal is not changed compared to an emotion in the speech signal.

[0014] DE 10 2013 224 417 B3 discloses a hearing aid device with a frequency analysis device configured to determine a current fundamental frequency value of a speech signal for a time period of the speech signal. A statistical evaluation device is configured to determine an average fundamental frequency value of the speech signal over several time periods. The hearing aid device further comprises a fundamental frequency modifier configured to modify the current fundamental frequency value to a modified fundamental frequency value such that a difference or a quotient of the current fundamental frequency value to the average fundamental frequency value is changed according to a specific function. In this way, a frequency range within which the fundamental frequency value varies can be changed.The hearing aid device further comprises a speech signal generator configured to generate a fundamental frequency modified speech signal based on the modified fundamental frequency value.

[0015] The invention is based on the object of providing a particularly suitable method for operating a hearing aid. In particular, the aim is to provide improved speech intelligibility in conversation situations with regard to emotion recognition. The invention is further based on the object of providing a particularly suitable hearing aid for implementing the method.

[0016] With regard to the method, the problem is solved according to the invention with the features of claim 1 and with regard to the hearing aid with the features of claim 6. Advantageous embodiments and further developments are the subject of the dependent claims (subclaims).

[0017] The advantages and embodiments described with regard to the method can also be applied to the hearing aid, and vice versa. Where method steps are described below, advantageous embodiments for the hearing aid result, in particular, from the fact that the hearing aid is designed to perform one or more of these method steps.

[0018] The method according to the invention is intended and configured to operate a hearing aid. The hearing aid is configured to be worn on or in the ear of a hearing aid user (hearing aid wearer), wherein the method is carried out in particular when the hearing aid is worn on or in the ear of the hearing aid user. The hearing aid is, in particular, a hearing aid device that is intended and configured to compensate for a hearing loss of the hearing aid user, in particular to compensate for it in terms of signaling.

[0019] The hearing aid has a receiving unit for receiving speech information and converting it into a speech signal. The hearing aid also has a signal processing unit for processing the speech signal. The signal processing unit converts the speech signal into an output signal (output audio signal, processed signal), which can be output to a hearing aid user via an output unit.

[0020] A "speech signal" is understood here and below to mean, in particular, an acoustic or electrical signal capable of transmitting, storing, or processing oral or spoken (speech) information. Such a speech signal contains, in particular, information generated by the human voice and may include words, sentences, tones, or other vocal utterances. A speech signal can exist in various forms, including analog sound waves, digital audio data, or other electrical signals that encode or transmit speech information.

[0021] The receiving unit comprises, for example, an acoustoelectrical transducer that records acoustic sound signals from an environment and converts them into a digital input signal. The transducer is preferably embodied as a microphone. Additionally or alternatively, the receiving unit can comprise, for example, a transceiver (RF receiver, T-coil, etc.) for receiving wireless radio signals and generating a corresponding input signal therefrom.

[0022] The particularly digital or electrical speech signal is usually part of the received input signal (input audio signal). For example, the receiving unit has a voice activity detection (VAD) to separate the speech signal from the rest of the input signal. Voice activity detection is understood in particular to mean the (signal-based) detection of the presence or absence of human speech. In other words, the speech signal is preferably the (digital or electrical) signal output by a voice activity detection system.

[0023] Voice activity detection can, for example, be part of the receiving unit. Preferably, however, voice activity detection is part of the signal processing unit, with the receiving unit generating an input signal from the speech information, which comprises the speech signal and a residual signal, with the speech signal only being isolated or separated during signal processing.

[0024] The output signal is, in particular, an audio or sound signal generated by an electro-acoustic transducer (speaker) as the output unit. The output unit converts the electrical output signal into an acoustic audio or sound signal.

[0025] According to the process, speech information is received and converted into a speech signal. The speech signal may originate, for example, from a speech utterance spoken by a speaker near the hearing aid user and captured by the hearing aid's receiver unit. Alternatively, the utterance may also come from a radio signal, such as a Bluetooth signal or a mobile network signal (e.g., media streaming or telephone conversations), which is transmitted to the receiver unit via an external device, such as a smartphone.

[0026] The speech signal is converted or processed into the output signal by the signal processing unit. In other words, the speech signal is modified or changed by signal processing in the signal processing unit. "Signal processing" here and below refers in particular to the conversion, manipulation, or analysis of the speech signal using digital or analog techniques so that the speech or speech information (or its acoustic information) in the output signal is understandable to the hearing aid user. Signal processing includes, among other things, filtering, amplification, modulation, and demodulation of the speech signal or the speech information it contains.

[0027] According to the method, at least one emotion-correlated feature is then determined from each of the speech signal and the output signal. In other words, at least one emotion-correlated feature is determined from each of the original (unprocessed) speech signal and the signal-processed speech signal (output signal).

[0028] An "emotion-correlated feature" is understood here and below to mean, in particular, a measurable or detectable property or parameter that is related to the speaker's emotional state, i.e., the source of the speech information, or that is correlated with such a state. The emotion-correlated feature can, for example, be a voice pitch or frequency. The emotion-correlated feature is determined, for example, based on a temporal, spectro-temporal analysis of the speech signal or output signal.

[0029] In particular, the emotion-correlated feature is a GeMAPS (Geneva Minimalistic Acoustic Parameter Set) feature or parameter, as described in F. Eyben et al., “The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for Voice Research and Affective Computing,” in IEEE Trans. Affective Computing, Vol. 7(2), July 2015.

[0030] The emotion-correlated feature is, for example, a pitch frequency, a jitter (i.e., a deviation from the pitch frequency), a first formant frequency, a second formant frequency, a third formant frequency, a first formant bandwidth, and so on.Formants (formant range, formant stretch), a difference in the peak amplitudes of consecutive pitch frequency periods, a loudness in the auditory spectrum, a harmonic-to-noise ratio, a summed energy between 50 Hz (Hertz) and 1000 Hz, a summed energy between 1 kHz (Kilohertz) and 5 kHz, a ratio between the strongest power peak between 0 kHz and 2 kHz and 2 kHz to 5 kHz, a spectral slope within 0 Hz to 500 Hz, a spectral slope in the range 500 Hz to 1500 Hz, a relative energy between the first, second and third formant, a power ratio between the peak of the spectral harmonics in the first, second and third formants and the power of the spectral peak at the pitch frequency, a power ratio between the pitch frequency and the second multiple of the pitch frequency, a power ratio between the pitch frequency and the highest harmonic in the third formant range.

[0031] In particular, the same parameter is determined as an emotion-correlated feature in both the speech signal and the output signal, so that the determined features of the speech signal and the output signal are comparable with each other. In a next step, the emotion-correlated features determined from the speech signal and the output signal are compared with each other, and a comparison result is determined based on the comparison. By comparing the results with the unprocessed speech signal, it is possible to assess the extent to which the emotion-correlated features or information of the speech signal have been altered by the signal processing in the output signal.

[0032] According to the invention, the signal processing unit or its signal processing is adjusted depending on the comparison result. According to the method according to the invention, the sum of the hearing aid algorithms is controlled and / or regulated in such a way that the acoustic features relevant for emotion recognition are not damaged (or are even improved), so that the hearing aid user is able to recognize emotions from the processed speech signal (i.e., in the output signal). The invention thus aims to analyze the input or speech signal with regard to its acoustic features relevant for emotion recognition and to control or limit the effect of each individual and all signal processing algorithms in such a way as to ensure that the acoustic features for emotion recognition are preserved as completely as possible.

[0033] According to the method, the speech signal is processed with a view to improving speech perception while simultaneously (numerically) preserving the emotion-related features contained in the speech information. In contrast to previous signal processing methods, the focus is on the conversation partner's emotions and their information retention.

[0034] For example, it is possible for a number of emotion-correlated features, i.e. several, at least two emotion-correlated features, to be determined from the signals. Which specific parameters or features are determined is initially irrelevant. For example, it is conceivable that different emotion-correlated features are determined and compared with each other for different signal processing of the speech signal. Which emotion-correlated features are influenced by which signal processing can be determined, for example, from past speech data or from corresponding experiments or trials. Different listening situations and environmental conditions or application scenarios may result in different relevant emotion-correlated features (e.g., conversation situations in loud or quiet environments, conversation situations in films, etc.).

[0035] Preferably, the set of parameters or features is always the same, regardless of the speech utterance. Different characteristics of the acoustic features characterize the emotions. For a given speech signal with signal processing enabled (e.g., denoising), a different characteristic is expected than when no signal processing is enabled. By comparing the features, the change can thus be used as an indicator to adjust the signal processing.

[0036] For example, with denoising (noise reduction) enabled, a feature that extracts information from a specific frequency band might be less pronounced than without denoising. In such a case, it can be assumed that the denoising algorithm is set too "aggressively" in the frequency range under consideration. According to the invention, in such a case, the denoising is set to a more "defensive" level. In particular, the denoising is adjusted such that the denoising (still) functions acceptably and the emotional feature remains pronounced.Such a trade-off between preserving emotional characteristics and speech intelligibility can be situation-dependent. For example, if the hearing aid wearer is standing on a train platform and wants to hear the loudspeaker announcement, preserved emotions in the hearing aid signal are less important and the focus is mainly or exclusively on speech intelligibility.

[0037] It is therefore conceivable, for example, that an acoustic environmental situation is recorded and classified using a corresponding classifier, whereby a weighting is made between the preservation of the emotional characteristics and speech intelligibility on the basis of the classified environmental situation.

[0038] Furthermore, it is conceivable that emotional characteristics attenuated by signal processing are artificially added back to the output signal.

[0039] In an advantageous embodiment, the comparison result is compared with at least one stored threshold, and the signal processing is adjusted depending on the threshold comparison. The speech signal is thus analyzed for its acoustic features relevant for emotion recognition. Thresholds or rules are defined for how the emotion-correlated features may be modified by the signal processing, i.e., the sum of the algorithms, so that the effect of each individual and all signal processing steps is controlled and limited. This ensures that the acoustic features and information in the output signal are preserved as much as possible.

[0040] Different thresholds are possible for different emotion-correlated features. Furthermore, different thresholds can be assigned to different signal processing functions, such as different algorithms, so that only those signal processing functions of the signal processing unit are adjusted or changed for which the threshold comparison yields a corresponding comparison result. This enables particularly effective and targeted (readjustment) or adjustment of the signal processing unit.

[0041] In one conceivable embodiment, the comparison result is determined as a deviation (deviation measure) between the emotion-correlated features of the speech signal and the output signal. The comparison result or deviation can, for example, be the square of a Euclidean distance of the form γ=∑i=1j|Fi−Fi'|2 where γ is the comparison result or deviation, F is the emotion-correlated feature of the speech signal, F' is the emotion-correlated feature of the output signal, i is a running index, and j is the number of used / relevant features. With a number of emotion-correlated features, the features thus form a feature vector, with the comparison result being, for example, the square of a Euclidean norm between the feature vectors of the speech signal and the output signal.

[0042] For example, if GeMAPS parameters are used as emotion-correlated parameters, j is equal to eighteen (18) if all eighteen GeMAPS parameters are used to examine or assess the deviation. In other words, feature vectors with eighteen entries are created for the speech signal and the output signal.

[0043] Preferably, the signal processing is adjusted to minimize the deviation. In other words, the or each emotion-correlated feature in the output signal is adapted as closely as possible to the corresponding emotion-correlated feature of the speech signal. This ensures that the information about the speaker's emotions is not distorted or is only minimally distorted. In other words, an output signal is generated that essentially corresponds to the speech signal with regard to the emotion information it contains.

[0044] The above formula represents a generic method for calculating the deviation. Depending on the specific application, the distance measure can be modified. For example, it is conceivable to supplement the formula by introducing weighting factors in order to weight some features more highly than others. This approach is particularly useful if the goal of the application is to capture a specific emotion, such as anger. In this example, the features that characterize "anger" would be given greater weight, thus placing particular emphasis on these features when minimizing the distance, thus preserving this emotion.

[0045] In an additional or alternative embodiment, the signal processing is adjusted based on the comparison result such that the at least one emotion-correlated feature in the output signal is amplified. In this embodiment, the method is used not only to retain acoustic features for emotion recognition in the output signal, but also to enhance them, i.e., to amplify them relative to the speech signal. This improves emotion recognition for the hearing aid user, which is particularly advantageous for hearing aid users whose ability to recognize emotions is impaired, e.g., due to hearing loss.

[0046] To ensure that emotional information is amplified and not artificially distorted, i.e. to ensure that the amplification only highlights emotions that are also contained in the original speech information, a preliminary check is carried out to determine whether the speech signal contains emotional information.

[0047] For example, the speech signal is analyzed using a probabilistic model that outputs a probability value for each emotion, indicating whether the emotion is present in the speech signal or not. Amplification of an emotion or an emotion-correlated feature is performed in particular when the probability of an emotion lies above a specified threshold. The threshold value is initially irrelevant. This can be determined, for example, from past conversation data or from corresponding experiments or trials. Such a probabilistic model is easy to derive. For example, with a suitable choice of emotion classification algorithm, the distance of the sample under consideration to the decision boundary can be used as a probability.

[0048] For example, during the comparison, the Euclidean norms of the specific feature vectors for the speech signal are compared with the output signal, and the signal processing of the signal processing unit is adjusted if a condition of the form F12+F22+⋯+Fj2≥F1'2+F2'2+⋯+Fj'2 is met, where F is the emotion-correlated feature of the speech signal, F' is the emotion-correlated feature of the output signal, i is a running index, and j is the number of used / relevant features. During the comparison, the lengths of the feature vectors are compared with each other, with the comparison result indicating which feature vector is longer, i.e., which signal has the greater information content with regard to emotion information. Signal processing is stopped if the feature vector of the output signal is less than or equal to the feature vector of the speech signal.

[0049] The fundamental problem here is that such a "live" modification of the emotional content of a speech utterance lacks the reference (the clear speech signal). For a comparison according to the above formula, for example, an estimate is made as to whether the analyzed recorded signal contains an emotion, and once a certain degree of certainty is reached, the emotion is artificially extracted. The latter can be achieved using emotion-specific prototype filters, which, for example, boost certain frequency ranges.

[0050] "Estimation" or "estimate" here and below refers to an approximate determination of the emotional information contained in the speech signal by evaluating the speech signal, for example, through pre-characterized measurements, stored tables or characteristic curves, or by means of statistical-mathematical methods. For example, a (robust) classifier or feature extractor is used to estimate whether the speech signal contains an emotion or an emotion-correlated feature.

[0051] In a useful further development, the length of the output signal feature vector is scaled, for example, with a stored scaling factor α (where α > 1), whereby the scaling factor is selected such that a desired amplification of the emotion-relevant features in the output signal is effected.

[0052] For example, it is conceivable to switch between the threshold comparison described above and the amplification of the emotion-correlated features depending on the listening situation. For example, in a quiet environment with little background noise, the threshold comparison is performed to reproduce the emotional information of the speech information as authentically as possible. In a noisy environment with a lot of background noise, for example, the emotion-correlated features are amplified so that they are more easily perceived by the hearing aid user.

[0053] The hearing aid according to the invention is intended and configured to be worn on or in the ear of a hearing aid user. The hearing aid is preferably further intended and configured to compensate for a hearing loss of the hearing aid user. In other words, in the preferred application, the hearing aid is a hearing aid designed to provide hearing aids for people with hearing impairments. The hearing aid is thus designed, in particular, as a hearing aid device.

[0054] The hearing aid according to the invention is intended for, and is suitable and configured for, carrying out a method described above. The hearing aid comprises a receiving unit for receiving speech information and converting it into a speech signal, a signal processing unit for converting a speech signal into an output signal, an output unit for outputting an output signal to a hearing aid user, a feature extractor for determining at least one emotion-correlated feature from the speech signal and from the output signal, and a controller (i.e., a control unit) for carrying out the method described above.

[0055] The controller is signal-linked to at least the signal processing unit and the feature extractor. The signal processing unit and the feature extractor can also be part of the controller. In other words, the signal processing and feature extraction functionalities can be integrated into the controller through programming and / or circuitry. In this case, the controller is also signal-linked directly to the receiving unit and the output unit.

[0056] The controller is generally configured—in terms of programming and / or circuitry—to implement the above-described inventive method. The controller is thus specifically configured to process the speech signal into the output signal by means of the receiving unit using signal processing and to determine at least one emotion-correlated feature from the speech signal and the output signal using the feature extractor. The controller is further configured to compare the determined features with one another and to adjust the signal processing based on the resulting comparison result.

[0057] In a preferred embodiment, the controller is formed, at least in its core, by a microcontroller with a processor and a data memory, in which the functionality for carrying out the method according to the invention is implemented in the form of operating software (firmware), so that the method - possibly in interaction with a device user - is carried out automatically when the operating software is executed in the microcontroller. Within the scope of the invention, however, the controller can alternatively also be formed by a non-programmable electronic component, such as an application-specific integrated circuit (ASIC) or an FPGA (Field Programmable Gate Array), in which the functionality for carrying out the method according to the invention is implemented using circuitry.

[0058] The hearing aid is designed to receive sound signals from the environment and output them to the hearing aid user. The hearing aid has a (hearing) device housing, which houses, for example, the receiving unit, the signal processing unit, the feature extractor, the controller, and the output unit. The device housing is designed so that it can be worn by the hearing aid user on the head and near the ear, e.g., in the ear, on the ear, or behind the ear.

[0059] The hearing aid has at least one acoustoelectrical input transducer, in particular a microphone, which is part of the receiving unit. During operation of the hearing aid, the input transducer receives sound signals (noises, tones, speech, etc.) from the environment and converts them into an electrical input signal (acoustic data, input audio signal). As part of the receiving unit, the hearing aid further comprises a speech recognition or voice activity detection (VAD) system, for example, as part of the signal processing unit, which generates a speech signal from the input signal or the acoustic data. The hearing aid further comprises, in particular, an electroacoustic output transducer, for example, in the form of a (miniature) loudspeaker, for generating an acoustic output signal based on an audio signal (output signal) generated by the signal processing unit. The output transducer can be part of the output unit.

[0060] The hearing aid according to the invention thus enables control and / or regulation of the sum of the hearing aid algorithms in such a way that the acoustic features relevant for emotion recognition are not damaged (or are even improved), so that the hearing aid user is able to recognize emotions from processed speech signals (i.e. in the output signal).

[0061] Exemplary embodiments of the invention are explained in more detail below with reference to a drawing. The single figure shows a simplified, schematic block diagram of a hearing aid.

[0062] The figure shows a hearing aid 2 designed specifically as a hearing aid.

[0063] The hearing aid 2 has a receiving unit 4 with at least one acousto-electrical input transducer 6, for example a microphone, and with a speech recognition unit 8 (voice activity unit).

[0064] During operation of the hearing aid 2, the input transducer 6 receives sound signals (noises, tones, speech, etc.) from the environment and converts them into an electrical input signal E. The speech recognition unit 8 is designed, for example, as part of a controller 10 and, during operation, generates a speech signal S from the input signal E if the received sound signals contain speech information.

[0065] The speech signal S is processed into an output signal A by a signal processing unit 12 of the controller 10. The signal processing of the signal processing unit 12 is the sum of algorithms for modifying the speech signal S into the output signal A, which are intended to improve speech comprehension for a hearing aid user, for example, by compensating for a hearing loss of the hearing aid user through signal processing.

[0066] The unprocessed speech signal S and the processed speech signal, i.e., the output signal A, are each fed to a feature extractor 14, which is, for example, also integrated into the controller 10. The feature extractor 14 determines a feature vector from the speech signal A and the output signal A, each with a number of emotion-correlated features as vector entries. For example, feature vectors with eighteen (18) vector entries or emotion-correlated features are determined, with the features of the speech signal S being designated F1, ..., F 18 and the characteristics of the output signal A with F1', ..., F 18 ' are marked.

[0067] The specific characteristics F1 to F 18 and F1' to F 18 ' are fed to a comparison unit 16 of the controller 10. The comparison unit 16 performs a feature comparison of the features F1 to F 18 and F1' to F 18' and produces a comparison result γ.

[0068] Depending on the comparison result γ, either a control and / or regulating signal R is generated, by means of which the signal processing of the signal processing unit 12, in particular its algorithms, is adjusted, or the output signal A is forwarded to an output unit 18. The output unit 18 is designed, for example, as an electro-acoustic output transducer, for example in the form of a miniature loudspeaker, and converts the output signal A into an acoustic sound signal for the hearing aid user.

[0069] In one possible embodiment, a square of a Euclidean distance between the feature vectors is determined as the feature comparison. The comparison result γ, or the deviation between the feature vectors, is thus γ=∑i=118|Fi−Fi'|2

[0070] The comparison result γ is compared with a stored threshold value, wherein the control and / or regulating signal R is generated if the comparison result γ is higher than the threshold value, i.e. if the deviation or the distance between the feature vectors is greater than the threshold value, and wherein the output signal A is output by means of the output unit 18 if the comparison result γ is smaller than the threshold value, i.e. if the deviation or the distance between the feature vectors is smaller than the threshold value.

[0071] The threshold is preferably determined based on audiological optimization. The threshold is, in particular, dimensioned such that, if the threshold is undershot, the emotional content in the output signal A is not significantly altered compared to the unprocessed speech signal S.

[0072] Using the control and / or regulation signal R, the gains during signal processing are adjusted, for example. The adjustment strategy depends largely on the signal processing algorithms used. How these algorithms specifically need to be modified to reduce the comparison result accordingly can be determined, for example, through appropriate testing. The described control and / or regulation mechanism is used particularly in the acoustic optimization of a hearing aid.

[0073] The claimed invention is not limited to the exemplary embodiments described above. Rather, other variants of the invention can also be derived therefrom by those skilled in the art within the scope of the disclosed claims without departing from the subject matter of the claimed invention. In particular, all individual features described in connection with the various exemplary embodiments can also be combined in other ways within the scope of the disclosed claims without departing from the subject matter of the claimed invention.

[0074] In an embodiment not shown, it is conceivable, for example, that the signal processing of the signal processing unit 12 is adjusted on the basis of the comparison result γ in such a way that the at least one emotion-correlated feature F1' to F 18 ' is amplified in the output signal A. List of reference symbols 2 hearing aids 4 Receiver unit 6 input converters 8 Speech recognition unit 10 controllers 12 Signal processing unit 14 Feature extractor 16 Comparison unit 18 Output unit E input signal S voice signal A output signal F1, ..., F 18 Features F1', ..., F 18 ' Features R Control and / or regulation signal

Claims

[1] Method for operating a hearing aid (2), comprising - a receiving unit (4) for receiving speech information and converting it into a speech signal (S), - a signal processing unit (12) for processing the speech signal (S) and generating an output signal (A), and - an output unit (18) for outputting an output signal (A) to a hearing aid user, - wherein speech information is converted into a speech signal (S), - wherein the speech signal (S) is processed by the signal processing unit (12) into an output signal (A) for output to the hearing aid user, - wherein at least one emotion-correlated feature (F1, ..., F 18 , F1', ..., F 18 ') is determined, - where the emotion-correlated features (F1, ..., F 18, F1', ..., F 18 ') are compared with each other and a comparison result is determined, and - wherein the signal processing unit (12) is adjusted depending on the comparison result. [2] Method according to claim 1, characterized by that the comparison result is compared with a stored threshold value, and that the signal processing unit (12) is adjusted depending on the threshold value comparison. [3] Method according to claim 1 or 2, characterized by , - that the comparison result is a deviation between the emotion-correlated features (F1, ..., F 18 ) of the speech signal (S) and the emotion-correlated features (F1', ..., F 18 ') of the output signal (A), and - that the signal processing unit (12) is adjusted such that the deviation is minimized. [4] Method according to one of claims 1 to 3, characterized bythat the signal processing unit (12) is set such that the at least one emotion-correlated feature (F1', ..., F 18 ') is amplified in the output signal (A). [5] Method according to one of claims 1 to 4, characterized by that the square of a Euclidean norm is used as the comparison result. [6] Hearing aid (2), comprising - a receiving unit (4) for receiving speech information and converting it into a speech signal (S), - a signal processing unit (12) for signal processing of the speech signal (S) and generating an output signal (A), and - an output unit (18) for outputting an output signal (A) to a hearing aid user, - a feature extractor (14) for determining at least one emotion-correlated feature (F1, ..., F 18 , F1', ..., F 18 ') from the speech signal (S) and from the output signal (A), - a comparison unit (16) for comparing the emotion-correlated features (F1, ..., F 18 , F1', ..., F 18 ') to determine a comparison result, and - a controller (10) for carrying out a method according to one of claims 1 to 5.

Citation Information

Patent Citations

  • Hearing aid device with fundamental frequency modification, method for processing a speech signal and computer program with program code for carrying out the method

    DE102013224417B3