Method for evaluating the speech quality of a speech signal using a hearing instrument
By using acousto-electric input converter and signal processing technology in listening devices, the pronunciation and pronunciation characteristics of speech signals are quantitatively collected, and the problem of difficulty in evaluating speech signal quality in the prior art is solved, and objective evaluation and improvement of speech signal quality is achieved.
Patent Information
- Application Number
- CN202110993782.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-28
- Filing Date
- 2021-08-27
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2041-08-27
AI Technical Summary
Prior art It is difficult to objectively evaluate the quality of speech signals when using hearing devices, especially when noise suppression algorithms may reduce sound quality.
By receiving the sound of the surrounding environment with the acousto-electrical input converter of the hearing device and converting it into an input audio signal, the pronunciation and pronunciation characteristics of the speech signal are quantitatively collected using signal processing technology to derive a quantitative measure of speech quality.
An objective evaluation of the quality of the voice signal is realized, the sound quality reduction caused by the noise suppression algorithm is avoided, and the intelligibility of the voice signal is improved.
Smart Images

Figure CN114121040B_ABST
Abstract
Description
Field of the Invention
[0001] The present invention relates to a method for evaluating the speech quality of a speech signal by means of a hearing device, wherein a sound containing the speech signal is received from the surroundings of the hearing device by means of an electroacoustic input transducer of the hearing device and converted into an input audio signal, and at least one characteristic of the speech signal is quantitatively acquired by analyzing the input audio signal by means of signal processing. Background Art
[0002] An important task when using a hearing device, such as a hearing aid, a headset or a communication device, is usually to output the speech signal to the user of the hearing device as precisely as possible, i.e., in particular, acoustically as intelligibly as possible. To this end, interference noise from the sound is usually suppressed in the audio signal generated from the sound having the speech signal, so as to emphasize the signal component representing the speech signal and thus improve its intelligibility. However, the sound quality of the generated output signal may generally be reduced by the algorithm for noise suppression, wherein artifacts may particularly be formed by the signal processing of the audio signal, and / or the auditory sensation generally feels less natural.
[0003] In most cases, noise suppression is implemented based on characteristic parameters, which primarily relate to the noise or the total signal, i.e., for example, the signal-to-noise ratio ("SNR"), the noise floor, or the level of the audio signal. However, this scheme for controlling noise suppression may ultimately result in the application of noise suppression even when there is obvious interference noise but there is no need to apply noise suppression at all because the speech component is still easily intelligible despite the interference noise. In this case, there is a risk of deterioration of the sound quality due to artifacts of the noise suppression, for example, without real necessity. On the contrary, in the case where the pronunciation of the speaker is very weak, a speech signal that is only superimposed with a small amount of noise and thus has a good SNR for the relevant audio signal may also have a low speech quality.
[0004] This can be avoided if, in the hearing device, the algorithm for noise suppression is controlled in special cases and the signal processing is controlled in general based on the quality of the speech signal component in the audio signal to be processed. However, for this purpose, it is necessary to be able to measure and acquire this quality completely. Summary of the Invention
[0005] Accordingly, the technical problem to be solved by the present invention is to provide a method by which the quality of the speech component in an audio signal to be processed by a hearing device can be objectively evaluated. In addition, the technical problem to be solved by the present invention is to provide a hearing device which is designed to objectively evaluate the quality of the speech component contained therein for an internal audio signal.
[0006] According to the present invention, the first-mentioned technical problem is solved by a method for evaluating the speech quality of a speech signal by means of a hearing device, wherein, by means of an electroacoustic input transducer of the hearing device, a sound containing the speech signal is received from the surroundings of the hearing device and converted into an input audio signal, wherein, by means of an analysis of the input audio signal by means of signal processing, in particular the signal processing of the hearing device and / or an auxiliary device connectable to the hearing device, at least one articulatory and / or prosodic characteristic of the speech signal is quantitatively acquired, and wherein a quantitative measure of the speech quality is derived on the basis of at least one articulatory or prosodic characteristic. Advantageous and partly inventive design solutions are the subject of the present invention and the following description.
[0007] According to the present invention, the second-mentioned technical problem is solved by a hearing device which comprises an electroacoustic input transducer and a signal processing means, in particular having a signal processor, wherein the electroacoustic input transducer is designed to receive a sound from the surroundings of the hearing device and convert it into an input audio signal, and wherein the signal processing means is designed to quantitatively acquire at least one articulatory and / or prosodic characteristic of the speech component contained in the input audio signal by means of an analysis of the input audio signal, and to derive a quantitative measure of the speech quality on the basis of at least one articulatory or prosodic characteristic.
[0008] The hearing device according to the present invention has the advantages of the method according to the present invention, which can in particular be implemented by means of the hearing device according to the present invention. The advantages mentioned below for the method and its extensions can be transferred mutatis mutandis to the hearing device according to the meaning.
[0009] The electroacoustic input transducer hereby particularly comprises any of the following transducers which are designed to generate an electrical audio signal from the sound of the surroundings, so that the air movement and the air pressure fluctuations caused by the sound are reproduced at the location of the transducer by corresponding oscillations of electrical parameters, in particular the voltage, in the generated audio signal. In particular, the electroacoustic input transducer can be provided by a microphone.
[0010] Signal processing is carried out especially with the aid of a corresponding signal processing device, which is designed to carry out calculations and / or algorithms for signal processing with at least one signal processor. Here, the signal processing device is especially arranged on a hearing device. However, the signal processing device can also be arranged on an auxiliary device designed to be connected to the hearing device for data exchange, such as a smartphone, a smartwatch, etc. The hearing device can, for example, transmit an input audio signal to the auxiliary device and carry out an analysis with the aid of the computing resources provided by the auxiliary device. Finally, as a result of the analysis, a quantitative measure can be transmitted back to the hearing device.
[0011] The analysis can be carried out directly on the input audio signal here, or on a signal derived from the input audio signal. This can especially be provided by an isolated speech signal component here; but it can also be provided by an audio signal, such as can be generated, for example, in a hearing device by a feedback loop with a compensation signal for compensating acoustic feedback, etc.; or by a directional signal generated from another input audio signal of another input transducer.
[0012] Here, the articulatory characteristics of the speech signal especially include formants, especially the accuracy of vowels, and the dominance of consonants, especially fricatives and / or plosives. It can be stated here that the higher the accuracy of the formants or the dominance and / or accuracy of the consonants, the higher the speech quality is set. The prosodic characteristics of the speech signal especially include the temporal stability of the fundamental frequency of the speech signal and the relative sound intensity of the stress.
[0013] Sound generation generally includes three physical components of a sound source: a mechanical oscillator, such as a string or a membrane, which sets the air around the oscillator into vibration; excitation of the oscillator (e.g., by plucking or stroking); and a resonator. By the excitation, the oscillator is set into oscillation, so that the air around the oscillator is in a pressure vibration through the vibration of the oscillator, and the pressure vibration propagates as a sound wave. Here, in a mechanical oscillator, mostly not only vibrations of a single frequency are excited, but also vibrations of different frequencies, where the spectral composition of the propagated vibrations determines the sound spectrogram. The frequencies of specific vibrations are usually provided here as integer multiples of a fundamental frequency and are called "harmonics" or overtones of this fundamental frequency. However, more complex spectral patterns can also be constructed, so that not all generated frequencies can be represented as harmonics of the same fundamental frequency. Here, the resonance of the frequencies generated in the resonance space is also related to the sound spectrogram, because specific frequencies generated by the oscillator in the resonance space usually decay relative to the main frequency of the sound.
[0014] When applied to the human voice, this means that the mechanical oscillator is provided by the vocal cords and their excitation in the air flowing from the lungs through the vocal cords, where the resonance space is mainly formed by the pharyngeal cavity and the oral cavity. The fundamental frequency of male voices is usually in the range of 60 Hz to 150 Hz, and that of female voices is usually in the range of 150 Hz to 300 Hz. Due to anatomical differences among individuals, not only regarding their vocal cords, but especially regarding the pharyngeal cavity and the oral cavity, different voices are first formed. By changing the volume and geometry of the oral cavity through the corresponding movements of the lower jaw and lips, the resonance space can be changed as follows, namely, formants, which are frequencies characteristic of the production of vowels, are formed. For each vowel, these formants are in an unchangeable frequency range (the so-called "formant range"), where the vowel can be clearly distinguished from other sounds by the first two formants F1 and F2 of a series of usually four formants (see "vowel triangle" and "vowel trapezoid"). Here, the formants are formed independently of the fundamental frequency, i.e., the frequency of the fundamental vibration.
[0015] In this sense, the accuracy of the formants is especially understood as the degree of concentration of sound energy on the mutually defined formant ranges, especially at the respective frequencies within the formant ranges, and the resulting determinability of each vowel according to the formants.
[0016] To produce consonants, the air stream flowing through the vocal cords is partially or completely blocked at at least one location, and thereby turbulence of the air stream is also formed. Therefore, only some consonants can be associated with a formant structure as clear as that of vowels, while other consonants have a broader-band frequency structure. However, consonants can also be associated with specific frequency bands in which the sound energy is concentrated. Due to the impulsive "noise nature" of consonants, these frequency bands are usually higher than the formant range of vowels, i.e., mainly in the range of approximately 2 to 8 kHz, while the ranges of the most important formants F1 and F2 of vowels usually end at approximately 1.5 kHz (F1) or 4 kHz (F2). Here, the accuracy of consonants is especially determined by the degree of concentration of sound energy on the corresponding frequency ranges and the resulting determinability of each consonant.
[0017] However, the distinguishability of the individual components of the speech signal and thus the possibility of resolving these components depend not only on the pronunciation aspect. The pronunciation aspect mainly relates to the acoustic accuracy of the smallest isolated sound events of speech, the so-called phonemes. The prosodic aspect also determines the speech quality, because special meanings can be given to a statement through intonation and stress, especially on several segments, i.e., several phonemes or groups of phonemes. For example, a question can be clarified by raising the pitch at the end of a sentence, or different meanings can be distinguished by stressing specific syllables in a word (see "um fahr en" and " um"(in Fahrenheit)” or by stressing a word to emphasize it. In this regard, the speech quality of a speech signal can also be quantitatively acquired according to prosodic characteristics, in particular as described above, by, for example, determining a measure of the pitch of the sound, i.e., the amount of change over time of its fundamental frequency, and a measure of the clarity of the contrast between the maximum amplitude and / or the maximum level.
[0018] Accordingly, a quantitative measure of speech quality can be derived from one or more of the aforementioned and / or additional quantitatively acquired articulatory and / or prosodic characteristics of the speech signal.
[0019] Preferably herein, as an articulatory characteristic of the speech signal, a characteristic parameter related to the accuracy of a preset formant of a vowel in the speech signal; a characteristic parameter related to the dominance of a consonant, in particular a fricative, in the speech signal; and / or a characteristic parameter related to the accuracy of the transition between voiced and voiceless sounds is acquired. The quantitative measure of speech quality can be provided directly by the aforementioned acquired characteristic parameters respectively, or formed according to the characteristic parameters, for example, by weighting two characteristic parameters of different formants, etc., or also by weighting at least two different characteristic parameters among the aforementioned characteristic parameters with each other, i.e., by forming a weighted average. The quantitative measure of speech quality is related to the speech generation of a speaker who may have defects (such as slurring or indistinctness) or even speech errors in a speech that is perceived as "clean", which correspondingly reduces the speech quality.
[0020] Differently from parameters related to the propagation of speech in an environment, such as the Speech Intelligibility Index ("SII") that weights each speech component and noise component in a frequency band manner, or the Speech Transmission Index ("STI") that acquires the influence of a transmission channel on the modulation depth by means of a test signal simulating the modulation of human speech, the current measure is in particular independent of the external characteristics of the transmission channel, such as propagation in a possibly reverberant space or a noisy environment, but preferably related only to the inherent characteristics of the speech generated by the speaker.
[0021] This particularly means that in a quiet environment and / or an environment with only low background noise, (based on a reference value that is preferably determined for a speech quality perceived as "very good") a reduced speech quality is recognized.
[0022] Advantageously, in order to collect characteristic parameters related to the dominance of consonants in the speech signal, a first energy contained in a low frequency range is calculated, a second energy contained in a higher frequency range above the low frequency range is calculated, and a relevant characteristic parameter is formed based on the ratio of the first energy and the second energy and / or the ratio weighted on the corresponding bandwidths of the mentioned frequency ranges. In particular, the speech signal can be smoothed over time in advance. In order to calculate the first and second energies, the input audio signal can be divided into a low and a higher frequency range, for example, by means of a filter bank and, if necessary, by means of a corresponding selection of the individual generated frequency bands. Preferably, the low frequency range is selected such that it lies within the frequency interval [0 Hz, 2.5 kHz], particularly preferably within the frequency interval [0 Hz, 2 kHz]. The higher frequency range is preferably selected such that it lies within the frequency interval [3 kHz, 10 kHz], particularly preferably within the frequency interval [4 Hz, 8 kHz].
[0023] It has been further confirmed to be advantageous that, in order to collect characteristic parameters related to the accuracy of the transition between voiced and unvoiced sounds, the voiced time series and the unvoiced time series are distinguished based on correlation measurement and / or based on the zero-crossing rate of the input audio signal or a signal derived from the input audio signal, the transition from the voiced time series to the unvoiced time series or from the unvoiced time series to the voiced time series is determined, the energy contained in the voiced or unvoiced time series before the transition is determined for at least one frequency range, and the energy contained in the unvoiced or voiced time series after the transition is determined for at least one frequency range, and the characteristic parameter is determined based on the energy before the transition and the energy after the transition.
[0024] This particularly means that: First, the voiced and unvoiced time series of the speech signal in the input audio signal are determined, and thus the transition from voiced to unvoiced or from unvoiced to voiced is identified. For at least one frequency range, which is preset according to empirical knowledge for the accuracy of the transition, the energy before the transition in the frequency range of the input audio signal or a signal derived therefrom is now determined. For example, this energy can be obtained from the voiced or unvoiced time series shortly before the transition. The energy in the relevant frequency range after the transition is determined, for example, by the unvoiced or voiced time series after the transition.
[0025] Based on these two energies, the eigenvalue can now be determined, and the eigenvalue can in particular enable an explanation of the change in the energy distribution during the transition. The eigenvalue can be determined, for example, as the quotient or relative deviation of the two energies before and after the transition. However, the eigenvalue can also be formed as a comparison of the energy before or after the transition with the overall (broadband) signal energy. However, in particular, the energies can also be determined separately before and after the transition for additional frequency ranges, so that the eigenvalue can additionally be determined based on the energies before and after the transition in the additional frequency bands, for example as the rate of change of the energy distribution into the participating frequency ranges during the transition (i.e., the comparison of the energy distribution in two frequency ranges before the transition with the distribution after the transition).
[0026] Based on the eigenvalue, a characteristic parameter related to the accuracy of the transition for the measure of speech quality can then be determined. For this purpose, the eigenvalue can be used directly, or the eigenvalue can be compared with a reference value determined in advance for good pronunciation, in particular based on the corresponding empirical knowledge (for example as a quotient or relative deviation). Usually, specific design solutions, in particular specific design solutions regarding the frequency ranges to be used and the boundary values or reference values, can be implemented based on the empirical results regarding the corresponding effectiveness of the corresponding frequency bands or groups of frequency bands. As at least one frequency range, in particular, the frequency bands 13 to 14, preferably 16 to 23, of the Bark scale can be used here. As an additional frequency range, in particular, a frequency range of lower frequencies can be used.
[0027] Preferably, in order to collect characteristic parameters related to the accuracy of a preset formant of a vowel in a speech signal, the acoustic energy (or a parameter related to the energy) of the speech signal concentrated in at least two different formant ranges is compared with each other. Particularly preferably, a signal component of the speech signal in at least one formant range in the frequency space is determined, a signal parameter related to the level is determined for the signal component of the speech signal in at least one formant range, and the characteristic parameter is determined according to the maximum value of the signal parameter related to the level and / or according to its temporal stability. In particular, here, as at least one formant range, the frequency range of the first formant F1 (preferably 250 Hz to 1 kHz, particularly preferably 300 Hz to 750 Hz) or the second formant F2 (preferably 500 Hz to 3.5 kHz, particularly preferably 600 Hz to 2.5 kHz) can be selected, or two formant ranges of the first and second formants can be selected. In particular, multiple first and / or second formant ranges associated with different vowels (i.e., the frequency ranges associated with the first or second formants of the corresponding vowels) can also be selected. Now, the signal component is determined for one or more selected formant ranges, and the signal parameter related to the level of the corresponding signal component is determined. Here, the signal parameter can be provided by the level itself or also by the maximum signal amplitude that is appropriately smoothed if necessary. Based on the temporal stability of the signal parameter (which is determined by the variance of the signal parameter within an appropriate time window) and / or based on the deviation of the signal parameter from its maximum value within an appropriate time window, now a statement about the accuracy of the formant can be made, i.e., a small variance and a small deviation from the maximum level of the pronounced sound (in particular, the length of the time window can be selected according to the length of the pronounced sound) represent high accuracy.
[0028] Advantageously, the fundamental frequency of the speech signal is collected in a time-resolved manner, and the characteristic parameter characterizing the temporal stability of the fundamental frequency is determined as the prosodic characteristic of the speech signal. For example, the characteristic parameter can be determined according to the relative deviation of the fundamental frequency accumulated over time, or by collecting multiple maximum and minimum values of the fundamental frequency within a preset time period. The temporal stability of the fundamental frequency is particularly important for the speech melody and the monotonicity of stress, and therefore, the quantitative collection also allows a statement about the speech quality of the speech signal.
[0029] Preferably, for a speech signal, in particular by analyzing the corresponding input audio signal or a signal derived therefrom, volume-related parameters, in particular amplitude and / or level, are acquired in a time-resolved manner, wherein a quotient of a maximum value of the volume-related parameters and an average value of the parameters determined within a preset time period is formed, and wherein, as a prosodic characteristic of the speech signal, a characteristic parameter is determined based on the quotient, the quotient being formed by the maximum value and the average value of the volume-related parameters within the preset time period. In this way, the definition of stress can be explained based on the indirectly acquired volume dynamics of the speech signal.
[0030] In an advantageous design, at least two characteristic parameters respectively characterizing pronunciation and / or prosodic characteristics are determined based on the analysis of the input audio signal, wherein a quantitative measure of the speech quality is formed based on the product of these characteristic parameters and / or based on the weighted average value and / or the maximum or minimum value of these characteristic parameters. This is particularly advantageous when a single measure of the speech quality is required or desired, or when a single measure of all pronunciation characteristics or all prosodic characteristics is desired.
[0031] Preferably, speech activity is detected and / or the SNR in the input audio signal is determined before acquiring at least one pronunciation and / or prosodic characteristic of the speech signal, and wherein the analysis of at least one pronunciation and / or prosodic characteristic of the speech signal is carried out based on the detected speech activity or the determined SNR. Thereby, the analysis of the speech quality of the speech signal can be restricted to cases where the speech signal actually exists or the SNR is particularly higher than a preset threshold value, so that it can be assumed that only when the signal component of the speech signal in the input audio signal can be recognized well enough can the corresponding evaluation be carried out. On the contrary, although the poor speech quality in the case of weak pronunciation and / or in the case of small prosodic features, such as weak stress, benefits from improvements by means of signal processing, in conventional signal processing, for a sufficiently high SNR, usually no measures are taken to emphasize or process the speech signal in a similar way.
[0032] Preferably, the hearing device is designed as a hearing aid. Here, the hearing aid can be provided by a monaural device or by a binaural device with two local devices, which are worn by the user of the hearing aid on their right or left ear, respectively. In addition to the input transducer mentioned, the hearing aid can in particular also have at least one further electroacoustic input transducer, which converts the sounds of the surrounding environment into a corresponding further input audio signal, so that a quantitative acquisition of at least one pronunciation and / or prosody characteristic of the speech signal can be carried out by analyzing a plurality of participating input audio signals. In the case of a binaural device, two of the input audio signals used can be generated in different local units of the hearing aid (i.e., on the left and right ears, respectively). Here, the signal processing device can in particular include signal processors of two local units, where preferably, the locally generated measure of speech quality is normalized in a suitable manner according to the pronunciation and / or prosody characteristic considered, by means of an average value or a maximum or minimum value formed for the two local units. Description of the Drawings
[0033] Embodiments of the present invention will subsequently be described in detail with reference to the drawings. Here, schematically:
[0034] Figure 1 A hearing aid is shown in a circuit diagram, which acquires sounds with speech signals; and
[0035] Figure 2 A method for determining a quantitative measure of a speech signal according to Figure 1 is shown in a block diagram.
[0036] Corresponding components and parameters have the same reference numerals in all the drawings, respectively. Detailed Description of the Embodiments
[0037] Figure 1 A hearing device 1 is schematically shown in a circuit diagram, which is currently designed as a hearing aid 2. The hearing aid 2 has an electroacoustic input transducer 4, which is designed to convert the sounds 6 of the surrounding environment of the hearing aid 2 into an input audio signal 8. A design of the hearing aid 2 with a further input transducer (not shown) that generates a corresponding further input audio signal from the sounds 6 of the surrounding environment is also conceivable here. The hearing aid 2 is currently constructed as a separate monaural device. It is also conceivable that the hearing aid 2 is designed as a binaural hearing aid with two local devices (not shown), which are worn by the user of the hearing aid 2 on their right and left ears, respectively.
[0038] The input audio signal 8 is fed to the signal processing device 10 of the hearing aid 2, where the input audio signal 8 is processed in particular according to the hearing requirements of the user of the hearing aid 2 and is amplified and / or compressed, for example, in a frequency band manner. For this purpose, the signal processing device 10 is designed in particular with the aid of a corresponding signal processor (not shown in detail in Figure 1 and a main memory addressable by the signal processor. The possible preprocessing of the input audio signal 8, such as the A / D conversion and / or preamplification of the generated input audio signal 8, should be regarded as part of the input converter 4 here.
[0039] The signal processing device 10 generates an output audio signal 12 by processing the input audio signal 8, and the output audio signal is converted into the output sound signal 16 of the hearing aid 2 by means of an electroacoustic output converter 14. Here, the input converter 4 is preferably provided by a microphone, and the output converter 14 is provided, for example, by a loudspeaker (such as a Balanced Metal Case Receiver), but the output converter 14 can also be provided by a bone conduction earphone, etc.
[0040] The sound 6 collected by the input converter 4 from the surroundings of the hearing aid 2 also includes a speech signal 18 of a loudspeaker not shown in detail and additional sound components 20. The additional sound components can in particular include directional and / or diffuse interference noise (interference sound or background noise), but can also include sounds that can be regarded as useful signals depending on the situation, such as music and acoustic warning signals or indication signals related to the surroundings.
[0041] The signal processing implemented in the signal processing device 10 for generating the output audio signal 12 can in particular include the suppression of signal components that suppress the interference noise contained in the sound 6, and the relative enhancement of the signal components representing the speech signal 18 relative to the signal components representing the additional sound components 20. Here, frequency-dependent or broadband dynamic compression and / or amplification and noise suppression algorithms can also be used in particular.
[0042] In order to be able to hear as well as possible the signal components representing the speech signal 18 in the input audio signal 8 in the output audio signal 12 and still be able to convey as natural an auditory sensation as possible to the user of the hearing aid 2 in the output sound 16, a quantitative measure of the speech quality of the speech signal 18 should be determined in the signal processing device 10 to control the algorithms applied to the input audio signal 8. According to Figure 2 Describe this.
[0043] Figure 2 It is shown in block diagram form for Figure 2Processing of the input audio signal 8 of the hearing aid 2. First, recognition of voice activity VAD is carried out for the input audio signal 8. If there is no notable voice activity (path "n"), then signal processing of the input audio signal 8 is carried out according to the first algorithm 25 to generate an output audio signal 12. The first algorithm 25 evaluates signal parameters of the input audio signal 8, such as level, background noise, transients, etc., in a broadband and / or especially band-wise manner in a pre-set manner in advance, and thereby determines various parameters applicable to the input audio signal 8, such as band-wise amplification factors and / or compression characteristic data (i.e., mainly inflection points, ratios, attack, release).
[0044] The first algorithm 25 can especially also be set to classify the hearing situation implemented in the sound 6, and set various parameters according to the classification, if necessary as a hearing program specifically set for the specific hearing situation. In addition, the individual hearing requirements of the user of the hearing aid 2 can also be considered for the first algorithm 25, so that the hearing impairment of the user can be compensated as well as possible by applying the first algorithm 25 to the input audio signal 8.
[0045] However, if notable voice activity is determined during the recognition of voice activity VAD (path "y"), then the SNR is next determined and compared with a preset boundary value Th SNR If the SNR is not higher than the boundary value, i.e., SNR ≤ Th SNR , then the first algorithm 25 is applied to the input audio signal 8 again to generate an output audio signal 12. However, if the SNR is higher than the preset boundary value Th SNR , i.e., SNR > Th SNR , then a quantitative measure 30 of the voice quality of the voice component 18 contained in the input audio signal 8 is determined in the manner described below for further processing of the input audio signal 8. For this purpose, the pronunciation and / or prosody characteristics of the voice signal 18 are quantitatively collected. The term voice signal component 26 contained in the input audio signal 8 is hereby understood as the signal component of the voice component 18 of the input audio signal 8 representing the sound 6, from which the input audio signal 8 is generated by means of the input converter 4.
[0046] To determine the mentioned quantitative measure 30, the input audio signal 8 is divided into respective signal paths.
[0047] For the first signal path 32 of the input audio signal 8, the central wavelength λ C is first determined and compared with a predetermined boundary value Th of the central wavelength λ If according to the above-mentioned boundary value Th of the central wavelength λIf it is determined that the signal components in the input audio signal 8 are of a sufficiently high frequency, then in the first signal path 32, after optionally appropriately selected time-based smoothing (not shown), signal components are selected for the low frequency range NF and the higher frequency range HF located above the low frequency range NF. A possible division could be, for example, that the low frequency range NF includes all frequencies f N ≤ 2500 Hz, in particular f N ≤ 2000 Hz, and the higher frequency range HF includes the frequencies f H : 2500 Hz < f H ≤ 10000 Hz, in particular 4000 Hz ≤ f H ≤ 8000 Hz or 2500 Hz < f H ≤ 5000 Hz.
[0048] The selection can be carried out directly in the input audio signal 8 or can also be implemented as follows, i.e., the input audio signal 8 is divided into individual frequency bands by means of a filter bank (not shown), where each frequency band is associated with the low or higher frequency range NF or HF according to the corresponding band boundaries.
[0049] Subsequently, a first energy E1 is determined for the signals contained in the low frequency range NF, and a second energy E2 is determined for the signals contained in the higher frequency range HF. Now, a quotient QE is formed with the second energy as the numerator and the first energy E1 as the denominator. In the case of appropriately selected lower and higher frequency ranges LF, HF, the quotient QE can now be considered as a characteristic parameter 33, which is related to the dominance of consonants in the speech signal 18. Thus, the characteristic parameter 33 enables an indication of the pronunciation characteristics of the speech signal component 26 in the input audio signal 8. For example, for a value of the quotient QE >> 1 (i.e., QE > Th QE , where the preset boundary value Th not shown in detail QE >> 1), a high dominance of consonants can be deduced, while for a value QE < 1, a low dominance can be deduced.
[0050] In the second signal path 34, in the input audio signal 8, a distinction 36 between the voiced time series V and the unvoiced time series UV is carried out based on correlation measurement and / or based on the zero-crossing rate of the input audio signal 8. Based on the voiced and unvoiced time series V or UV, a transition TS from the voiced time series V to the unvoiced time series UV is determined. The length of the voiced or unvoiced time series can be, for example, between 10 and 80 ms, in particular between 20 and 50 ms.
[0051] Now for at least one frequency range (e.g., a suitably determined selection of particularly effective frequency bands, such as bands 16 to 23 on the Bark scale, or bands 1 to 15 on the Bark scale), the energy Ev of the voiced time series V before the transition TS and the energy En of the unvoiced time series UV after the transition TS are determined separately. In particular, for more than one frequency range, the corresponding energies before and after the transition TS can also be determined separately. Now, for example, by the relative change ΔE TS or by the quotient of the energies Ev, En before and after the transition TS (not shown) to determine how the energy changes at the transition TS.
[0052] The measure of the energy change, i.e., the current relative change, is now compared with a boundary value Th determined in advance for good pronunciation for the energy distribution at the transition. E In particular, the characteristic parameter 35 can be formed according to the relative change ΔE TS in comparison with the boundary value Th E or according to the ratio of the relative change ΔE TS to this boundary value Th E or according to the relative deviation of the relative change ΔE
[0053] from this boundary value Th. The characteristic parameter 35 is related to the pronunciation of the transition between voiced and unvoiced sounds in the speech signal 18 and can thus enable an explanation of additional pronunciation characteristics of the speech signal component 26 in the input audio signal 8. Here, the following statement generally applies, i.e., in the frequency range related to voiced and unvoiced sounds, the faster the change in energy distribution occurs, i.e., the more time - definably it occurs, the more precisely the transition between the voiced and unvoiced time series is pronounced.
[0054] However, for the characteristic parameter 35, the energy distribution can also be considered in two frequency ranges (e.g., according to the above - mentioned frequency ranges on the Bark scale, or in lower and higher frequency ranges NF, HF) by, for example, the quotient of the corresponding energies or comparable characteristic values, and the characteristic parameter takes into account the change in the quotient or characteristic value at the transition. Thus, for example, the rate of change of the quotient or characteristic parameter can be determined and compared with a reference value determined in advance for the rate of change.
[0054] To form the characteristic parameter 35, the transition from the unvoiced time series can also be observed in a similar way. Specific design solutions, especially specific design solutions regarding the frequency ranges to be used and the boundary values or reference values, can generally be achieved based on empirical results regarding the corresponding effectiveness of the respective frequency bands or groups of frequency bands.
[0055] In the third signal path 38, the fundamental frequency f of the speech signal component 26 in the input audio signal 8 is acquired in a time - resolved manner, and the variance of the fundamental frequency f G is determined for the fundamental frequency f G and based on the variance of the fundamental frequency f GDetermine the temporal stability 40. The temporal stability 40 can be used as a characteristic parameter 41, which enables an account of the prosodic characteristics of the speech signal component 26 in the input audio signal 8. Here, a larger variance of the fundamental frequency f G can be considered as an indicator of better speech intelligibility, while a monotonic fundamental frequency f G has less speech intelligibility.
[0056] In the fourth signal path 42, the level LVL is acquired in a time-resolved manner for the input audio signal 8 and / or for the speech signal component 26 contained therein, and a time average MN is formed in a time period 44 that is preset in particular according to the corresponding empirical knowledge LVL . Furthermore, the maximum value MX of the level LVL is determined within the time period 44 LVL . Now, the maximum value MX of the level LVL LVL is divided by the time average MN of the level LVL LVL , and thus a characteristic parameter 45 related to the volume of the speech signal 18 is determined, which enables a further account of the prosodic characteristics of the speech signal component 26 in the input audio signal 8. Instead of the level LVL, another parameter related to the volume and / or energy content of the speech signal component 26 can also be used here.
[0057] The characteristic parameters 33, 35, 41, or 45 respectively determined as described in the first to fourth signal paths 32, 34, 38, 42 can now be considered separately as quantitative measures 30 of the quality of the speech component 18 contained in the input audio signal 8. Based on this measure, the second algorithm 46 is now applied to the input audio signal 8 for signal processing. Here, the second algorithm 46 can be generated from the first algorithm 25 through corresponding changes in one or more parameters of the signal processing that are implemented based on the relevant quantitative measure 30, or the second algorithm can be set up as a completely independent hearing program.
[0058] In particular, a single value can also be determined as a quantitative measure 30 of the speech quality based on the characteristic parameters 33, 35, 41, or 45 determined as described, for example, through the product or weighted average of the characteristic parameters 33, 35, 41, 45 (schematically shown by combining the characteristic parameters 33, 35, 41, 45 in Figure 2 ). The weighting of the individual characteristic parameters can be carried out in particular according to a weighting factor determined in advance based on experience, and the weighting factor can be determined according to the effectiveness of the articulatory or prosodic characteristics of the speech quality captured by the corresponding characteristic parameters.
[0059] Although the present invention has been described in detail by way of preferred embodiments, the present invention is not limited to the disclosed examples, and other variants can be derived therefrom by those skilled in the art without departing from the scope of protection of the present invention.
[0060] List of reference numerals
[0061] 1 Hearing device
[0062] 2 Hearing aid
[0063] 4 Input transducer
[0064] 6 Sound of the surrounding environment
[0065] 8 Input audio signal
[0066] 10 Signal processing device
[0067] 12 Output audio signal
[0068] 14 Output transducer
[0069] 16 Output sound
[0070] 18 Speech signal
[0071] 20 Sound component
[0072] 25 First algorithm
[0073] 26 Speech signal component
[0074] 30 Quantitative measure of speech quality
[0075] 32 First signal path
[0076] 33 Characteristic parameter
[0077] 34 Second signal path
[0078] 35 Characteristic parameter
[0079] 36 Discrimination
[0080] 38 Third signal path
[0081] 40 Temporal stability
[0082] 41 Characteristic parameter
[0083] 42 Fourth signal path
[0084] 44 Time period
[0085] 45 Characteristic parameter
[0086] 46 Second algorithm
[0087] ΔE TS Relative change in energy (upon transition)
[0088] λ C Central wavelength
[0089] E1 First energy
[0090] E2 Second energy
[0091] Ev Energy (before transition)
[0092] En Energy (after transition)
[0093] f G Fundamental frequency
[0094] LVL Level
[0095] HF Higher frequency range
[0096] MN LVL Time average (of level)
[0097] MX LVL Maximum value of level
[0098] NF Low frequency range
[0099] QE Quotient
[0100] SNR Signal-to-noise ratio (SNR)
[0101] Th λ Boundary value (of central wavelength)
[0102] Th E Boundary value (of relative change in energy)
[0103] Th SNR Boundary value (of SNR)
[0104] TS Transition
[0105] V Voiced time series
[0106] VAD Voice activity detection
[0107] UV Unvoiced time series
Claims
1. A method for evaluating the speech quality of a speech signal (18) by means of a hearing instrument (1), - wherein sound (6) containing a speech signal (18) is received from the surroundings of the hearing device (1) by means of an acoustic-electric input converter (4) of the hearing device (1) and the sound is converted into an input audio signal (8), -in, At least one pronunciation characteristic of the speech signal (18) is quantitatively acquired by analyzing the input audio signal (8) by means of signal processing, and - wherein a quantitative measure (30) of the speech quality is derived from at least one pronunciation characteristic, Among them, as the pronunciation characteristics of the speech signal (18), the collection - a first characteristic variable associated with the accuracy of a predetermined formant of a vowel in the speech signal (18), and / or - a second characteristic variable (33) associated with the dominance of consonants in the speech signal (18), and / or a third characteristic parameter (35) related to the accuracy of the transition between voiced and unvoiced sounds, It is characterized in that In order to acquire a second characteristic variable (33) which is related to the dominance of consonants in the speech signal (18), - calculating a first energy (E1) contained in a low frequency range NF, wherein the low frequency range NF is selected within the frequency interval [0 Hz, 2.5 kHz], - calculating a second energy (E2) contained in a higher frequency range HF above the lower frequency range NF, wherein the higher frequency range HF is selected within the frequency interval [3 kHz, 10 kHz], and forming a second characteristic variable as a function of a ratio (QE) of the first energy (E1) to the second energy (E2) and / or a ratio weighted over the respective bandwidths of the mentioned lower frequency range NF, the higher frequency range HF, In order to acquire a third characteristic parameter (35) related to the accuracy of the transition between voiced and unvoiced sounds, - distinguishing (36) the voiced time series V from the unvoiced time series UV according to a correlation measure and / or according to the zero crossing rate, - determining a transition (TS) from a voiced time series V to an unvoiced time series UV or from an unvoiced time series UV to a voiced time series V, - determining for at least one frequency range the energy (Ev) contained in the voiced time sequence V or the unvoiced time sequence UV before the transition (TS), and determining for at least one frequency range the energy (En) contained in the unvoiced time sequence UV or the voiced time sequence V after the transition (TS), and - determining a third characteristic variable (35) as a function of the energy (Ev) before the transition (TS) and as a function of the energy (En) after the transition (TS), In order to acquire a first characteristic parameter related to the accuracy of a predetermined formant of a vowel in a speech signal (18), - determining a signal component of the speech signal (18) within the range of at least one formant in frequency space, - determining a level-dependent signal variable for a signal component of the speech signal (18) in at least one formant range, and - determining the first characteristic variable as a function of the maximum value and / or the temporal stability of the level-dependent signal variable.
2. The method according to claim 1, - wherein at least one prosodic characteristic of the speech signal (18) is further quantitatively acquired by analyzing the input audio signal (8) by means of signal processing, and -in, In addition, a quantitative measure (30) of the speech quality is determined as a function of at least one prosodic property of the speech signal (18).
3. The method according to claim 2, in, The fundamental frequency (f) of the speech signal (18) is acquired in a time-resolved manner. G ),and Among them, the fundamental frequency (f G ) is determined as a prosodic property of the speech signal (18).
4. The method according to claim 2 or 3, in, For speech signals (18), volume-related parameters (LVL) are acquired in a time-resolved manner. In which, within a preset time period (44), a maximum value (MX LVL ) and a mean value (MN) of the parameter (LVL) determined within a preset time period (44) LVL ) and In which, as a prosodic characteristic of the speech signal (18), a fifth characteristic parameter (45) is determined based on the quotient, which is a maximum value (MX) of a volume-related parameter (LVL) within a preset time period (44). LVL ) and mean value (MN LVL )form.
5. The method according to claim 1 or 2, in, determining at least two of the first characteristic parameter, the second characteristic parameter, the third characteristic parameter, the fourth characteristic parameter and the fifth characteristic parameter based on an analysis of the input audio signal (18), and Therein, a quantitative measure (30) for the speech quality is formed as a function of the product of the at least two characteristic variables and / or as a function of a weighted average value of the at least two characteristic variables.
6. The method according to claim 1 or 2, in, detecting voice activity (VAD) and / or determining a signal-to-noise ratio (SNR) in an input audio signal (18) before acquiring at least one pronunciation characteristic and / or prosodic characteristic of the speech signal, and In this case, an analysis is carried out with respect to at least one pronunciation property and / or prosodic property of the speech signal (18) as a function of the detected speech activity (VAD) or the determined signal-to-noise ratio (SNR).
7. A hearing device (1), comprising: - an acoustic-electrical input transducer (4) designed to receive sound (6) from the surroundings of the hearing device (1) and to convert the sound into an input audio signal (8), and - A signal processing device (10) which is designed to quantitatively detect at least one pronunciation characteristic of a component of a speech signal (18) contained in the input audio signal (8) based on an analysis of the input audio signal (8) and to derive a quantitative measure (30) of the speech quality based on the at least one pronunciation characteristic in accordance with a method according to any of the preceding claims.
8. The hearing device (1) according to claim 7, which is designed as a hearing aid (2).
Citation Information
Patent Citations
Auditory-articulatory analysis for speech quality assessment
US20040002852A1
Audio-based method, system, and apparatus for measurement of voice quality
US20040167774A1