Method for operating a hearing device
By integrating a receiving unit, a signal processing unit, and a feature extractor into a hearing device, emotion-related features in the speech signal are identified and optimized, solving the problem of insufficient emotion recognition in dialogue scenarios and achieving better speech comprehension and emotion recognition results.
Patent Information
- Application Number
- CN202510624580.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2025-05-15
- Publication Date
- 2025-12-02
AI Technical Summary
Existing hearing devices struggle to effectively identify a speaker's emotions in conversational situations, resulting in hearing-impaired wearers not receiving complete signal information and affecting the accuracy of emotion recognition.
By integrating a receiving unit, a signal processing unit, and a feature extractor into a hearing device, emotion-related features in the speech signal are determined. Through signal processing control algorithms, emotion recognition features are preserved or improved in the output signal. Speech activity recognition and signal processing technologies are used, combined with feature comparison and threshold control, to adjust signal processing to optimize emotion recognition.
It improves speech understanding and emotion recognition capabilities in conversational situations, ensuring that emotion-related features are not impaired or are improved during signal processing, thereby enhancing the emotion recognition capabilities of hearing device users.
Smart Images

Figure CN121056802A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for operating a hearing device. Furthermore, this invention relates to a hearing device for performing said method. Background Technology
[0002] Generally, electronic devices that assist the hearing of people wearing hearing aids are referred to as hearing devices. In particular, this invention relates to hearing devices configured to fully or partially compensate for the hearing loss of users with hearing impairments. Such hearing devices are also referred to as "hearing aids" (HA). Furthermore, there are hearing devices that protect or improve the hearing of users with normal hearing, for example, to improve speech comprehension in complex hearing situations. Such devices are also referred to as "personal sound amplification products" (PSAP). Finally, the term "hearing device" as used herein also includes headphones (wired or wireless, with or without active interference suppression), earphones, etc., worn on or in the ear, as well as implantable hearing devices, such as cochlear implants. Hearing devices can also be part of an AR (Augmented Reality) or VR (Virtual Reality) system, used to output acoustic information of virtual sound sources to the user.
[0003] Generally speaking, hearing devices, and more specifically, hearing aids, are typically designed to be worn on a user's head, particularly in or on the ear, especially as behind-the-ear (BTE) or in-the-ear (ITE) devices. Internally, hearing devices typically have at least one output transducer that converts the output audio signal fed in for output into a signal that the user can perceive as sound, and then outputs that signal to the user.
[0004] In most cases, the output converter is constructed as an electroacoustic converter, which converts the (electrical) output audio signal into airborne sound, which is then output to the user's ear canal. In the case of behind-the-ear hearing devices, the output converter, also known as a "receiver," is typically integrated into the housing of the hearing device outside the ear. In this case, the sound output by the output converter is guided into the user's ear canal via a sound tube. Alternatively, the output converter can also be positioned inside the ear canal, thus outside the housing worn behind the ear. This type of hearing device is also called an RIC device, following the English term "Receiver In Channel." Hearing devices worn in the ear (comprehensively in-canal) are also called CIC devices, as these devices are designed to be so small that they do not protrude out of the ear canal.
[0005] In other configurations, the output transducer can also be an electromechanical transducer that converts the output audio signal into solid-borne sound (vibration), which, for example, is output to the user's skull. Furthermore, there are implantable hearing devices, particularly cochlear implants, and hearing devices whose output transducers directly stimulate the user's auditory nerve.
[0006] In addition to the output converter, hearing devices often have at least one (acoustic-to-electrical) input converter. When the hearing device is in operation, the input converter, or each input converter, records airborne sound from the environment surrounding the hearing device and converts that airborne sound into an input audio signal (i.e., an electrical signal that conveys information about the ambient sound). This input audio signal (also called the "recorded sound signal") is typically output to the user in its raw or processed form, for example, to implement a so-called transparency mode in headphones, for active interference suppression, or (e.g., in hearing aids) to achieve improved noise perception for the user.
[0007] In addition, hearing devices often have a signal processing unit (signal processor). The signal processing unit processes the input audio signal, or each input audio signal, (i.e., modifies its sound information). The signal processing unit then outputs the corresponding processed audio signal (also referred to as the "output audio signal" or "modified sound signal") to an output converter and / or external device.
[0008] Hearing devices, for example, provide various additional (hearing / auditory device) functions within the scope of signal processing to improve auditory benefits for hearing aid wearers (HAWs). Examples of such functions include: Own Voice Detection (OVD), Voice Activity Detection (VAD), Active Noise Reduction (ANR), Active Occlusion Reduction (AOR); streaming audio information (e.g., music); recognition of various bodily signals (e.g., health status); and recognition of specific events and responses to them (e.g., fall detection, sending an alarm when the user falls). These functions are often implemented as processing algorithms within software or signal processing units.
[0009] In the application of hearing systems or devices, handling conversational situations is one of the core issues. This is primarily because users of hearing systems frequently glean important information from personal conversations. In these conversations, the emotions of the other person are also crucial information.
[0010] As is known from BW Schuller's "Speech emotion recognition: two decades in anutshell, benchmarks and ongoing trends" (Communications of the ACM, Volume 61, Issue 5, pp. 90-99), it is possible to identify and classify a speaker's emotions in a speech signal. For example, multiple related acoustic features exist, and these acoustic features can be extracted from the speech signal. Furthermore, text content or emotion keywords can also be used to assist acoustic (emotion) analysis.
[0011] To understand the acoustic expression patterns of different emotions and other emotional tendencies and processes, acoustic parameters can be extracted from speech signals. The underlying theoretical assumption here is that emotional processes alter the voluntary excitation and tension of the striatal muscles in different ways, thereby affecting the production of sound and speech at the level of phonation and articulation. These alterations can then be assessed using different parameters of the acoustic waveform.
[0012] When a hearing device is in operation, multiple algorithms work to program the signal in order to compensate for hearing loss, but also provide additional auxiliary algorithms for optimal speech comprehension and / or comfortable listening. All of these manipulations of the original input signal can affect acoustic features important for emotion recognition. That is, the hearing device's compensation for hearing loss may be accompanied by a deterioration in emotion recognition. Therefore, in conversational situations, the hearing device wearer may not be able to obtain all the signal information they need for emotion recognition. Summary of the Invention
[0013] The technical problem to be solved by the present invention is to provide a particularly suitable method for operating a hearing device. In particular, it is to provide improved speech understanding in dialogue situations, with enhanced emotion recognition. Furthermore, the technical problem to be solved by the present invention is to provide a particularly suitable hearing device for performing the method.
[0014] According to the present invention, the above-mentioned technical problems are solved by utilizing the features of the present invention in terms of method and in terms of hearing devices. Advantageous designs and extensions are the subject of the following description.
[0015] The advantages and design schemes listed regarding the methods described can also be applied to hearing devices, and vice versa. The following description of the method steps specifically illustrates how advantageous designs for hearing devices are derived by constructing them for implementing one or more of these method steps.
[0016] The method according to the invention is set and designed for operating a hearing device. Here, the hearing device is configured for wearing on or in the ear of a hearing device user (hearing device wearer), wherein, particularly here, the method is performed when the hearing device is worn on or in the ear of the hearing device user. The hearing device is particularly a hearing aid device, which is set and configured for compensating for hearing loss of the hearing device user, particularly by means of signal technology.
[0017] Here, the hearing device has a receiving unit for receiving speech information and converting it into a speech signal. Furthermore, the hearing device has a signal processing unit for processing the speech signal using signal technology (signal processing). Here, the signal processing unit converts the speech signal into an output signal (output audio signal, processed signal), which can be output to the hearing device user via an output unit.
[0018] "Speech signal" here and below should be understood in particular as an acoustic or electrical signal capable of transmitting, storing, or processing spoken or articulated (speech) information. Such speech signals specifically contain information produced by human voices, and this information can include words, sentences, tones, or other vocal expressions. Speech signals can exist in various forms, including as analog sound waves, digital audio data, or other electrical signals that encode or transmit speech information.
[0019] The receiving unit may include, for example, an acoustic-to-electric transceiver that records acoustic sound signals from the environment and converts them into digital input signals. Preferably, the transceiver is configured as a microphone. Additionally or alternatively, the receiving unit may include, for example, a transceiver (RF receiver, T-coil, etc.) for receiving wireless radio signals, thereby generating a corresponding input signal.
[0020] Here, in particular, digital or electrical voice signals are typically part of the received input signal (input audio signal). The receiving unit, for example, has Voice Activity Detection (VAD) for separating or distinguishing the voice signal from the remaining input signal. Here, Voice Activity Detection should be understood specifically as (through signal technology) identifying whether human voice is present or absent. In other words, preferably, the voice signal is a (digital or electrical) signal output by Voice Activity Detection.
[0021] Speech activity recognition can be, for example, part of a receiving unit. However, it is preferred that speech activity recognition be part of a signal processing unit, wherein the receiving unit generates an input signal from speech information, the input signal including a speech signal and a residual signal, wherein the speech signal is separated or isolated during signal processing.
[0022] The output signal is specifically an audio or sound signal generated by an electroacoustic transducer (loudspeaker) that serves as the output unit. In other words, the output unit converts the electrical output signal into an acoustic audio or sound signal.
[0023] According to the method, voice information is received and converted into a voice signal. Here, the voice signal originates, for example, from a speech produced by a speaker near the user of the hearing device and captured by the receiving unit of the hearing device. Alternatively, the speech may also originate from a radio signal transmitted to the receiving unit via an external auxiliary device, such as a smartphone, a Bluetooth signal, or a mobile radio network signal (e.g., a media stream or telephone conversation).
[0024] Here, the speech signal is converted or processed into an output signal by the signal processing unit. In other words, the speech signal is modified or altered through signal processing by the signal processing unit. "Signal processing" here and below should be understood in particular as the transformation, manipulation, or analysis of the speech signal using digital or analog techniques, so that the speech or speech information (or its acoustic information) in the output signal is understandable to the user of the hearing device. Signal processing includes, among other things, filtering, amplification, modulation, and demodulation of the speech signal or the speech information contained therein.
[0025] According to the method, at least one emotion-related feature is subsequently determined from the speech signal and the output signal. In other words, at least one emotion-related feature is determined from the raw (unprocessed) speech signal and the speech signal processed by signal technology (output signal).
[0026] "Emotion-related features" should be understood here and below, in particular, as measurable or acquireable characteristics or parameters that, in conjunction with the speaker's emotional state—that is, the source of the speech information—are correlated with that state, or vice versa. Emotion-related features may, for example, be frequency ranges or sound frequencies. For instance, emotion-related features can be determined using spectral-time analysis of the speech signal or output signal.
[0027] Emotion-related features are particularly those described in F. Eyben et al., “The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for Voice Research and Affective Computing”, IEEE Trans. Affective Computing, Vol. 7(2), Juli 2015, or GeMAPS features (GeMAPS: Geneva Minimalistic Acoustic Parameter Set).
[0028] Features related to emotion include: pitch frequency, jitter (i.e., deviation from the pitch frequency), frequency of the first formant, frequency of the second formant, frequency of the third formant, bandwidth of the first formant (formant range, formant distance), difference in peak amplitude of consecutive pitch frequency periods, loudness in the auditory spectrum, overtone-to-noise ratio, total energy between 50 Hz and 1000 Hz, total energy between 1 kHz and 5 kHz, ratio between the strongest power peaks between 0 kHz and 2 kHz and between 2 kHz and 5 kHz, spectral slope between 0 Hz and 500 Hz, spectral slope between 500 Hz and 1500 Hz, relative energy between the first, second, and third formants, power ratio between the peak power of spectral harmonics in the first, second, and third formants and the power of the peak power of the spectral frequency at the pitch frequency, power ratio between the pitch frequency and the second harmonic of the pitch frequency, and power ratio between the pitch frequency and the highest harmonic in the range of the third formant.
[0029] Here, in both the speech signal and the output signal, the same parameters are specifically identified as emotion-related features, allowing specific features of the speech signal and the output signal to be compared. In subsequent method steps, the emotion-related features determined from the speech signal and the output signal are compared, and the comparison result is determined. By comparing with the unprocessed speech signal, the degree to which the emotion-related features or information of the speech signal have been altered by signal processing in the output signal can be determined based on the comparison result.
[0030] According to the present invention, the signal processing unit or its signal processing is configured based on the comparison results. According to the method of the present invention, the sum of the algorithms of the hearing device is controlled and / or adjusted so that the acoustic features related to emotion recognition are not impaired (or even improved), thereby enabling the hearing device user to identify emotions from the processed speech signal (i.e., in the output signal). Therefore, the object of the present invention is to analyze the acoustic features of the input signal or speech signal related to emotion recognition, and to control or limit the effect of each individual algorithm and all algorithms of the signal processing, so as to ensure that the acoustic features used for emotion recognition are preserved as completely as possible.
[0031] Therefore, according to the method, signal processing of the speech signal not only improves speech perception but also simultaneously (numerically) preserves the emotion-related features contained in the speech information. Thus, compared to signal processing to date, the focus is on the other party's emotion and the preservation of their information.
[0032] Here, for example, several emotion-related features can be determined from the signal, i.e., multiple, or at least two emotion-related features. Specifically determining which parameters or features are important here. For example, it is conceivable to determine different emotion-related features for different signal processing of the speech signal and compare them with each other. Which emotion-related features are affected by which signal processing can be determined, for example, from past speech data or from corresponding experiments or tests. Different emotion-related features may be obtained in certain situations for different hearing conditions, environmental conditions, or application scenarios (e.g., dialogue in noisy or quiet environments, dialogue in movies, etc.).
[0033] Ideally, the parameter set or feature set should remain the same regardless of the speech expression. Here, different manifestations of acoustic features characterize emotions. For a given speech signal and the enabled signal processing (e.g., denoising), different feature manifestations are expected compared to when signal processing is disabled. Therefore, by comparing features, changes can be used as indicators to adjust signal processing.
[0034] For example, when denoising is enabled, it may occur that the features from which information is extracted in a particular frequency band have a smaller representation compared to when denoising is not performed. In this case, it can be assumed that the denoising algorithm is set too "aggressively" within the considered frequency range. According to the invention, denoising is set "defensively" in this case. In particular, denoising is set such that it (still) works in an acceptable manner, and emotional features remain prominent. This trade-off between preserving emotional features and speech understanding can be situational, such as when a hearing device wearer is standing on a platform and wants to hear a loudspeaker announcement, where the emotion preserved in the hearing device signal is less important, and the focus is primarily or solely on speech understanding.
[0035] In other words, one could think of collecting acoustic environmental conditions and classifying them using a corresponding classifier, where a weighted average is applied between preserving emotional features and speech understanding based on the classified environmental conditions.
[0036] Furthermore, it is conceivable to artificially add the emotional characteristics, which have been weakened by signal processing, back into the output signal.
[0037] In an advantageous implementation, the comparison result is compared with at least one stored threshold, and signal processing is configured based on this threshold comparison. That is, acoustic features of the speech signal relevant to emotion recognition are analyzed, wherein thresholds or rules are defined regarding how emotion-related features should be altered by the signal processing—that is, the sum of changes made by the algorithm—thereby controlling and limiting the effects of each individual and all signal processing. This ensures that acoustic features and information are preserved as much as possible in the output signal.
[0038] Different thresholds are possible for different emotion-related features. Furthermore, different signal processing methods, such as different algorithms, can be associated with different thresholds, thereby setting or changing the signal processing unit's threshold comparison to provide the corresponding comparison result. This allows for particularly effective and targeted (re)adjustment or setting of the signal processing unit.
[0039] In a conceivable design, the comparison result is used to determine, for example, the deviation (deviation metric) between the emotion-related features of the speech signal and the output signal. Here, the comparison result or deviation can be represented, for example, as the square of the Euclidean distance:
[0040]
[0041] Where γ is the comparison result or bias, F is the emotion-related feature of the speech signal, F' is the emotion-related feature of the output signal, i is the running index, and j is the number of features used / related. Therefore, for multiple emotion-related features, the features form a feature vector, where the comparison result is, for example, the squared Euclidean norm between the feature vectors of the speech signal and the output signal.
[0042] If, for example, the GeMAPS parameters are used as parameters related to emotion, then j equals eighteen (18), for example, when all eighteen GeMAPS parameters are used to check or judge deviation. In other words, for the speech signal and the output signal, for example, a feature vector with eighteen terms is set.
[0043] Here, signal processing is preferably configured in a manner that minimizes bias. In other words, the emotion-related features in the output signal, or each emotion-related feature, are matched as closely as possible to the corresponding emotion-related features in the speech signal. This ensures that information about the speaker's emotion is not distorted, or is distorted only to a minimum. In other words, an output signal is generated that substantially corresponds to the speech signal in terms of the emotional information it contains.
[0044] The preceding formula is a general method for calculating bias. The distance metric can be modified based on the specific application. For example, it is conceivable to supplement the formula by introducing weighting factors to give some features a higher weight compared to others. This approach is particularly meaningful when the application aims to preserve a specific emotion, such as anger. In this example, the feature representing "anger" will be given a stronger weight, thus focusing particularly on this feature when minimizing the distance, thereby paying special attention to the preservation of this emotion.
[0045] In an additional or alternative implementation, signal processing is configured based on the comparison result such that at least one emotion-related feature is amplified in the output signal. That is, in this implementation, the method is used not only to preserve the acoustic features for emotion recognition in the output signal, but also to improve, i.e., amplify, the acoustic features for emotion recognition relative to the speech signal. This improves emotion recognition for users of hearing devices, which is particularly advantageous for users of hearing devices whose ability to recognize emotions is adversely affected, for example, due to hearing loss.
[0046] To ensure that emotional information is amplified without artificially distorting it, i.e., to ensure that the amplification only highlights the emotions that are also contained in the original speech information, the speech signal is checked beforehand to see if it contains emotional information.
[0047] For example, a probabilistic model can be used to examine a speech signal, calculating the probability of the presence of each emotional output. Specifically, when the probability of an emotion exceeds a stored threshold, the emotion or emotion-related features are amplified. The size of this threshold is initially irrelevant; it can be determined, for example, from past dialogue data or corresponding trials or tests. This probabilistic model is easily derived. For instance, with an appropriately chosen emotion classification algorithm, the distance between the considered sample and the decision boundary can be used as the probability.
[0048] For example, during the comparison process, the Euclidean norm of a specific feature vector of the speech signal is compared with the output signal, and the signal processing of the signal processing unit is set when the following condition is met:
[0049]
[0050] Where F represents the emotion-related features of the speech signal, F' represents the emotion-related features of the output signal, i is the running index, and j is the number of features used / related. That is, during the comparison process, the lengths of the feature vectors are compared, and the comparison result indicates which feature vector is longer, i.e., which signal has greater information content in terms of emotional information. Here, signal processing is set when the feature vector of the output signal is less than or equal to the feature vector of the speech signal.
[0051] Here, in principle, there is a problem: in this "live" modification of the emotional content of spoken expression, there is no reference (i.e., a definite speech signal). This involves comparing the recorded signal according to the preceding formula, estimating whether it carries emotion, and artificially emphasizing that emotion with a certain degree of reliability. The latter can be achieved using prototype filters typical of emotions, such as those that enhance a specific frequency range.
[0052] "Estimation" or "probability" here and below should be understood as approximating the emotional information contained in a speech signal by evaluating the speech signal, for example, with the aid of pre-characterized measurements, stored tables or characteristic curves, or with the aid of mathematical statistical methods. For example, using a (robust) classifier or feature extractor to estimate whether a speech signal contains emotion or emotion-related features.
[0053] In a suitable scaling scheme, for example, the length of the output signal feature vector is scaled using a stored scaling factor α (where α > 1), wherein the scaling factor is chosen such that the emotion-related features in the output signal are amplified as desired.
[0054] Here, for example, it is conceivable to switch between the threshold comparison described above and the amplification of emotion-related features, depending on the hearing condition. For instance, in a quiet environment with little background noise, a threshold comparison is performed to reproduce the emotional information of the speech information as without distortion as possible, while in a noisy environment with much background noise, emotion-related features are amplified, for example, so that the hearing device user can more easily perceive the emotion-related features.
[0055] The hearing device according to the invention is configured and set for wearing on or in the ear of a hearing device user. Furthermore, the hearing device is preferably configured and set for compensating for hearing loss in a hearing device user. In other words, the hearing device is, in a preferred application, constructed as a hearing aid for the hearing-impaired. That is, the hearing device is particularly implemented as a hearing assistive device.
[0056] The hearing device according to the invention is configured and suitably adapted for performing the methods described above. The hearing device includes a receiving unit for receiving speech information and converting it into a speech signal, a signal processing unit for converting the speech signal into an output signal, an output unit for outputting the output signal to a user of the hearing device, a feature extractor for determining at least one emotion-related feature from the speech signal and correspondingly from the output signal, and a controller (i.e., a control unit) for performing the methods described above.
[0057] The controller is coupled to at least the signal processing unit and the feature extractor via signal technology. Here, the signal processing unit and the feature extractor can also be part of the controller. In other words, the functions of signal processing and feature extraction can be integrated into the controller through programming and / or circuitry techniques. Here, the controller is also directly coupled to the receiving unit and the output unit via signal technology.
[0058] Here, the controller is generally configured (through programming and / or circuitry techniques) to perform the method according to the invention described above. Specifically, the controller is configured to process a speech signal into an output signal via signal processing using a receiving unit, and to determine at least one emotion-related feature accordingly from the speech signal and the output signal using a feature extractor. Furthermore, the controller is configured to compare specific features with each other and to adjust the signal processing based on the resulting comparison.
[0059] In a preferred design, the controller is formed at its core by a microcontroller having a processor and data storage, wherein the functions for performing the method according to the invention are implemented in the form of operating software (firmware) through programming techniques, so that when the operating software is executed in the microcontroller, the method is executed automatically (interacting with the device user if necessary). However, within the scope of the invention, the controller can also alternatively be formed by non-programmable electronic components, such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), wherein the functions for performing the method according to the invention are implemented using circuit techniques.
[0060] Here, the hearing device is configured to record sound signals from the environment and output them to the hearing device user. The hearing device has a housing, in which components such as a receiving unit, a signal processing unit, a feature extractor, a controller, and an output unit are housed. The housing is configured so that the hearing device user can wear it on their head near the ear, for example, in the ear, on the ear, or behind the ear.
[0061] Hearing devices have at least one electroacoustic input transducer, particularly a microphone, which is part of the receiving unit. During operation, the input transducer records sound signals (noise, pitch, speech, etc.) from the environment and converts them into electrical input signals (acoustic data, input audio signals). Here, the hearing device, as part of the receiving unit, also incorporates speech recognition or voice activity recognition (VAD), for example, as part of a signal processing unit, which generates a speech signal from the input signal or acoustic data. Furthermore, hearing devices particularly have an electroacoustic output transducer, for example in the form of a (miniature) loudspeaker, for generating an acoustic output signal based on the audio signal (output signal) generated by the signal processing unit. Here, the output transducer may be part of the output unit.
[0062] Therefore, the hearing device according to the invention enables control and / or adjustment of the sum of the hearing device algorithms so as not to impair (or even improve) the acoustic features related to emotion recognition, thereby enabling the hearing device user to recognize emotions from the processed speech signal (i.e., in the output signal). Attached Figure Description
[0063] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0064] Here, Figure 1 A simplified schematic diagram of a hearing device is shown. Detailed Implementation
[0065] Figure 1 The hearing device 2, in particular, is implemented as a hearing aid.
[0066] The hearing device 2 has a receiving unit 4, which has at least one audio-to-electrical input converter 6, such as a microphone, and a voice activity unit 8.
[0067] When the hearing device 2 is running, the input converter 6 records sound signals (noise, pitch, speech, etc.) from the environment and converts them into an electrical input signal E. The speech recognition unit 8 is implemented, for example, as part of the controller 10, and generates a speech signal S from the input signal E during operation when the received sound signal contains speech information.
[0068] The speech signal S is processed into an output signal A by the signal processing unit 12 of the controller 10. Here, the signal processing of the signal processing unit 12 is the sum of algorithms used to modify the speech signal S into the output signal A. These algorithms should, for example, improve speech understanding for users of hearing devices by compensating for hearing loss with signal processing techniques.
[0069] The unprocessed speech signal S and the processed speech signal, i.e., the output signal A, are respectively fed to feature extractor 14, which is also integrated into controller 10, for example. Feature extractor 14 determines a feature vector from the speech signal S and the output signal A, the feature vector having multiple emotion-related features as vector terms. For example, a feature vector with eighteen (18) vector terms or emotion-related features is determined, wherein the features of the speech signal S are represented by F1, ..., F2. 18 The characteristics of the output signal A are represented by F1', ..., F1'. 18 'express.
[0070] Specific features F1 to F 18 and F1' to F 18 The data is fed to the comparison unit 16 of the controller 10. The comparison unit 16 compares features F1 to F... 18 and F1' to F 18 Perform feature comparison and generate comparison result γ.
[0071] Based on the comparison result γ, a control and / or adjustment signal R is generated. This control and / or adjustment signal R is used to set the signal processing of the signal processing unit 12, particularly its algorithm, or to forward the output signal A to the output unit 18. The output unit 18 is implemented, for example, as an electroacoustic output converter in the form of a miniature loudspeaker, and converts the output signal A into an acoustic sound signal for users of hearing devices.
[0072] In one possible implementation, as a feature comparison, the square of the Euclidean distance between the feature vectors is determined. The comparison result γ, or the deviation between the feature vectors, is therefore...
[0073]
[0074] Here, the comparison result γ is compared with the stored threshold. When the comparison result γ is greater than the threshold, that is, when the deviation or distance between the feature vectors is greater than the threshold, a control and / or adjustment signal R is generated. When the comparison result γ is less than the threshold, that is, when the deviation or distance between the feature vectors is less than the threshold, an output signal A is output by means of the output unit 18.
[0075] The threshold is preferably determined based on auditory optimization. In particular, the threshold value is determined such that, when it is below the threshold, the emotional content in the output signal A does not change significantly compared to the unprocessed speech signal S.
[0076] Here, for example, the amplification in signal processing is adjusted using a control and / or adjustment signal R, where the adjustment strategy depends largely on the algorithm used for signal processing. For example, it can be determined through corresponding experiments how the algorithm used for signal processing needs to be changed to correspondingly reduce the comparison result. The described control and / or adjustment mechanism is used particularly when optimizing hearing devices using sound technology.
[0077] The claimed invention is not limited to the embodiments described above. Rather, those skilled in the art can derive other variations of the invention within the scope of the disclosure without departing from the subject matter of the claimed invention. Furthermore, all the individual features described in particular in conjunction with different embodiments can be combined in other ways within the scope of the disclosure without departing from the subject matter of the claimed invention.
[0078] In one embodiment not shown, it is conceivable, for example, that the signal processing of the signal processing unit 12 is configured according to the comparison result γ, such that at least one emotion-related feature F1' to F1' is included in the output signal A. 18 'enlarge.
[0079] List of reference numerals
[0080] 2 Hearing devices
[0081] 4 receiving units
[0082] 6-input converter
[0083] 8 speech recognition units
[0084] 10 controllers
[0085] 12 signal processing units
[0086] 14 Feature Extractors
[0087] 16 Comparison Units
[0088] 18 output units
[0089] E input signal
[0090] S-voice signal
[0091] A output signal
[0092] F1、…、F 18 feature
[0093] F1'、…、F 18 'feature
[0094] R control and / or regulation signals
Claims
1. A method for operating a hearing device (2), the hearing device (2) comprising: - Receiving unit (4), used to receive voice information and convert it into voice signal (S). - Signal processing unit (12), used to perform signal processing on the speech signal (S) and generate an output signal (A), and - Output unit (18), used to output the output signal (A) to the user of the hearing device. - in, Convert speech information into speech signals (S). - The signal processing unit (12) processes the speech signal (S) into an output signal (A) so that the output signal (A) can be output to the user of the hearing device. - Wherein, at least one emotion-related feature (F1, ..., F2) is determined accordingly by the speech signal (S) and the output signal (A). 18 F1', ..., F 18 '), - Wherein, the emotion-related features (F1, ..., F2) determined by the speech signal (S) and the output signal (A) 18 F1', ..., F 18 Compare them with each other, and determine the comparison result. - The signal processing unit (12) is configured based on the comparison result.
2. The method according to claim 1, characterized in that, The comparison result is compared with a stored threshold, and the signal processing unit (12) is set according to the threshold comparison.
3. The method according to claim 1 or 2, characterized in that, - Determine the emotion-related features (F1, ..., F) of the speech signal (S). 18 The emotion-related features (F1', ..., F1') of the output signal (A) and the output signal (A) 18 The deviation between ') is used as the comparison result, and - Configure the signal processing unit (12) to minimize the deviation.
4. The method according to any one of claims 1 to 3, characterized in that, The signal processing unit (12) is configured such that at least one emotion-related feature (F1', ..., F1') is included in the output signal (A). 18 ')enlarge.
5. The method according to any one of claims 1 to 4, characterized in that, The square of the Euclidean norm is used as the comparison result.
6. A hearing device (2), said hearing device (2) comprising: - Receiving unit (4), used to receive voice information and convert it into voice signal (S). - Signal processing unit (12), configured to process the speech signal (S) using signal technology and generate an output signal (A), and - Output unit (18), used to output the output signal (A) to the user of the hearing device. - Feature extractor (14) for determining at least one emotion-related feature (F1, ..., F2) from the speech signal (S) and the output signal (A) respectively. 18 F1', ..., F 18 '), - Comparison unit (16) for comparing emotion-related features (F1, ..., F2) determined by the speech signal (S) and the output signal (A). 18 F1', ..., F 18 ') to compare and determine the comparison result, and - Controller (10) for performing the method according to any one of claims 1 to 5.