Voice detector built into the hearing device
The integration of a voice accelerometer with microphones in audio devices allows for reliable detection and control of user's voice, addressing the issue of simultaneous amplification of background noise and user's voice, enhancing audio quality by maintaining user voice at its original level.
Patent Information
- Application Number
- CN202080101014.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-05-29
AI Technical Summary
While existing listening devices enhance user voice and background noise, it is difficult to effectively distinguish and control the user's own voice, resulting in voice amplification that will affect the user experience.
Using a combination of voice accelerometer (VAC) signal and microphone signal, the user's voice is accurately detected by identifying the fundamental tone and cepspectral coefficient calculations, and the volume is adjusted using a noise suppressor and enhancer to ensure that the user's voice is at the original level.
It realizes reliable detection and control of user voice, reduces background noise interference, maintains user voice at the original level, adapts to the ambient volume, and improves the user experience of listening equipment.
Smart Images

Figure CN115668370B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wearable devices. In particular, the present invention relates to a voice detector for a hearing device and a voice detection method. In addition, the present invention also relates to the hearing device itself and a hearing system including a plurality of such hearing devices. Background Art
[0002] At the same time, wireless earphones or other wearable devices are also widely used as mobile accessories for electronic devices. Traditionally, wearable devices are used for listening to music (playing). When there is a microphone in the wearable device, the wearable device can also be used for telephone services, for example, in cooperation with an electronic device. Recently, there has also been an increasing interest in using wearable devices to listen to ambient sounds.
[0003] When a user uses a wearable device, in addition to a speaker disposed inside the user's ear, the wearable device may also include one or more microphones disposed outside the user's ear. When a user uses a wearable device, the wearable device itself is usually inserted into the user's ear. Natural side listening is usually required and external signals are reproduced so that the user can hear ambient sounds in a similar way as if the user were not wearing the wearable device at all. In this case, the enhanced hearing function in the wearable device includes a set of audio signal processing methods for improving the auditory effect to obtain clarity or pleasantness.
[0004] In addition, for hearing-impaired users, the enhanced hearing function means using a wearable device similar to a hearing aid. However, anyone can benefit from the enhanced hearing function because it can control the external voice level. For example, if one person speaks too loudly to another person, the other person can adjust the voice of that person to a tolerable level. Correspondingly, if someone speaks very quietly, a wearable device can be used to enhance that person's voice.
[0005] However, the problem at hand is that simple enhancement (i.e., enhancing all aspects) also enhances the user's own voice and background noise. In particular, the user's own voice may become too loud as a result.
[0006] Specifically, simple enhancement applies a gain to the microphone signal (possibly modified by natural side listening) and plays the generated modified microphone signal on the speaker of the wearable device. Usually, such enhancement also amplifies low-level background noise, which can then be reduced by using a noise suppressor. However, such enhancement also amplifies the user's own voice; or, in the case of negative gain, reduces the user's own voice.
[0007] For hearing aids, more complex methods involving the detection of the user's own voice have been studied. For example, one method uses two microphones and detects the user's own voice through an adaptive filter between the two microphone signals. Another method assumes implanting sensors in the user's head. Then, detection is performed by comparing signal strengths.
[0008] However, the first method requires two microphones and is not suitable for affordable devices with only one microphone. In the second method, the detection does not consider the typical characteristics of speech and may be affected by chewing, etc.
[0009] Therefore, there is a need to improve the voice detection of hearing devices themselves. Summary of the Invention
[0010] In view of the above problems and disadvantages, the present invention aims to improve the voice detection of hearing devices such as wearable devices. Therefore, the object of the present invention is to provide a voice detector for a hearing device, which can reliably and easily detect voice, especially the user's own voice of the hearing device. Specifically, the improved voice detector should overcome the above disadvantages.
[0011] In a first aspect, the present invention relates to a voice detector for a hearing device, the voice detector being configured to: acquire one or more microphone signals; acquire a voice accelerometer (VAC) signal; identify whether a fundamental tone exists in the VAC signal based on the one or more microphone signals; and if a fundamental tone is identified in the VAC signal, determine whether the fundamental tone is related to a voice signal.
[0012] For example, the voice accelerometer may be a low-noise, high-bandwidth, and time-division multiplexing (TDM) three-axis micro-electro-mechanical system (MEMS) accelerometer. Since the voice accelerometer has a high bandwidth, it is particularly suitable for wearable devices or smart earphones, in which the voice accelerometer can significantly improve the audio quality, especially in systems using MEMS microphones. The hearing device may be an in-ear headphone. The VAC signal may be a signal corresponding to the vibration caused by the wave propagation in the human body when the user of the hearing device speaks. The VAC can be used to pick up such vibrations and convert them into the VAC signal. Each of the microphone signals is a signal corresponding to the sound wave propagating in the air, picked up by one or more microphones and converted into the microphone signal. The VAC may not be affected by such sound waves propagating in the air, and thus can be specifically used to detect the user's own voice inside the ear.
[0013] The voice detector according to the first aspect has the following advantages, namely, it can reliably and easily detect the user's own voice. In addition, the voice detector also has the following advantages, namely, it can amplify or reduce the ambient sound, while the user's own voice can be maintained at (or at least close to) its original level. In addition, the voice detector also has the following advantages, namely, it can be used to adapt the user's voice volume to the ambient voice volume and the way the user hears his own voice respectively.
[0014] In one implementation of the voice detector according to the first aspect, the voice detector is used to: determine a first VAC threshold according to the one or more microphone signals; identify whether there is a pitch in the VAC signal according to the first VAC threshold.
[0015] The pitch can be calculated at a lower sampling rate (e.g., 2 kHz) only when the signal strength of the VAC signal is high enough (the first VAC threshold) in order to have medium complexity.
[0016] In addition, this implementation has the following advantages, namely, the pitch can be detected in a simple and reliable manner.
[0017] In another implementation of the voice detector according to the first aspect, the voice detector is further used to: determine whether the pitch is related to the voice signal according to the determined second VAC threshold.
[0018] This implementation has the following advantages, namely, the pitch can be detected in a simple and reliable manner, and thus the related voice can also be detected.
[0019] In another implementation of the voice detector according to the first aspect, the voice detector is further used to: additionally, if the frequency of the pitch is within a predefined frequency range, determine that the pitch is related to the voice signal.
[0020] This implementation has the following advantages, namely, the pitch can be detected in a simple and reliable manner only according to the predefined frequency range.
[0021] In another implementation of the voice detector according to the first aspect, the first VAC threshold is determined by comparing the signal power of the current frame of the one or more microphone signals with the average signal power of multiple frames of the one or more microphone signals.
[0022] Therefore, the current frame of each microphone signal can be compared with the average value of multiple frames of the same microphone signal. Alternatively, the average value of the current frames of multiple microphone signals can be compared with the average value of multiple frames of multiple microphone signals.
[0023] In another implementation of the voice detector according to the first aspect, if the signal power of the current frame of the one or more microphone signals is higher than the average signal power of the one or more microphone signals, the first VAC threshold has a higher value; and / or if the signal power of the current frame of the one or more microphone signals is equal to or lower than the average signal power of the one or more microphone signals, the first VAC threshold has a lower value.
[0024] In other words, if there is sufficient signal in the one or more microphones, the threshold for pitch detection is lower; if there is no signal in the one or more microphones, the threshold for pitch detection is higher. This is because only such own voice that can be heard in the one or more microphones will be attenuated.
[0025] In another implementation of the voice detector according to the first aspect, in order to identify whether there is a pitch in the VAC signal, the voice detector is configured to: determine whether the pitch detected in at least one frame of the VAC signal is a male pitch or a female pitch.
[0026] In another implementation of the voice detector according to the first aspect, in order to identify whether there is a pitch in the VAC signal, the voice detector is further configured to: search for the determined male pitch in other frames of the VAC signal; or search for the determined female pitch in other frames of the VAC signal.
[0027] This improves the accuracy of detecting the user's voice in the VAC signal because periodic sounds outside the male or female frequency range can be excluded.
[0028] In another implementation of the voice detector according to the first aspect, in order to identify whether there is a pitch in the VAC signal, the voice detector is further configured to: calculate one or more cepstral coefficients according to the VAC signal; calculate the pitch according to the one or more cepstral coefficients based on the first VAC threshold and the second VAC threshold.
[0029] This provides a simple and reliable method for detecting the pitch in the VAC signal, and the VA signal is likely to be related to voice.
[0030] In another implementation of the voice detector according to the first aspect, the one or more cepstral coefficients are calculated according to the following formula:
[0031] c = IFFT(abs(FFT(x))^2),
[0032] Wherein, x represents the VAC signal, FFT represents the Fast Fourier Transform, and IFFT represents the Inverse Fast Fourier Transform.
[0033] This has the advantage that the voice detector can calculate the cepstral coefficients in a computationally efficient manner.
[0034] In another implementation of the voice detector according to the first aspect, in order to identify whether the fundamental tone exists in the VAC signal, the voice detector is further configured to: determine the maximum cepstral coefficient corresponding to a certain frequency range; if the value obtained by dividing the maximum cepstral coefficient by the signal power is greater than the first VAC threshold, it is identified that the fundamental tone exists in the VAC signal.
[0035] In another implementation of the voice detector according to the first aspect, the voice detector is further configured to: if the normalized maximum cepstral coefficient of the current frame of the VAC signal is higher than the second VAC threshold, determine that the fundamental tone is related to the voice signal in the current frame of the VAC signal; wherein, the second VAC threshold is determined by the average signal power of multiple frames of the VAC signal.
[0036] This helps to avoid voice detection for very low fundamental tones (e.g., below 65 Hz), which are considered chewing or other non-voice activities inside the user's mouth.
[0037] According to a second aspect, the present invention relates to a hearing device, comprising: a voice detector according to one of the first aspect and its implementations; a noise suppressor for generating one or more modified microphone signals by selectively applying a gain to the voice signal in the one or more microphone signals; wherein, the voice signal corresponds to the voice signal in the VAC signal detected by the voice detector.
[0038] This has the advantage that noise such as background noise can be suppressed, or the microphone signal can be enhanced, while the voice signal can be kept unaffected. Of course, compared with the background noise or environmental noise in the microphone signal, the loudness of the voice signal can also be relatively increased or decreased. In addition, the hearing device can be computationally efficient.
[0039] In an implementation of the hearing device according to the second aspect, the noise suppressor is further configured to generate the one or more modified microphone signals by suppressing the background noise signal in the one or more microphone signals.
[0040] This has the advantage that unnecessary noise can be suppressed, making other sounds (e.g., music) or eavesdropping as well as the user's speech (the speech signal) easier to hear.
[0041] In another implementation of the hearing device according to the second aspect, the hearing device further comprises: an enhancer for applying enhancement to the one or more modified microphone signals, in particular enhancement determined according to user input.
[0042] This has the advantage that the microphone signal can be adjusted according to the user's preference. At the same time, the hearing device according to the second aspect can leave the user's own speech unaffected.
[0043] In another implementation of the hearing device according to the second aspect, the gain applied to the speech signal is determined according to the enhancement applied to the one or more modified microphone signals.
[0044] In other words, according to whether enhancement is applied to the one or more microphone signals and what kind of enhancement is applied, the gain can be determined and selectively applied to the speech signal. In this way, the speech signal can be adjusted relative to other sounds in the one or more microphone signals.
[0045] In another implementation of the hearing device according to the second aspect, the hearing device is further configured to: select the gain according to the enhancement such that the signal power of the speech signal in the one or more modified microphone signals is equal to the signal power of the speech signal in the one or more microphone signals.
[0046] In another implementation of the hearing device according to the second aspect, if no enhancement is applied, the gain is zero; if negative enhancement is applied, the gain is positive; if positive enhancement is applied, the gain is negative.
[0047] In another implementation of the hearing device according to the second aspect, the hearing device further comprises: one or more microphones for generating the one or more microphone signals; and / or a VAC for generating the VAC signal.
[0048] According to a third aspect, the present invention relates to a system, comprising: a first hearing device according to one of the second aspect and its implementations, the first hearing device comprising: a first voice detector according to one of the first aspect and its implementations, configured to obtain one or more first microphone signals; and comprising a first noise suppressor; a second hearing device according to one of the second aspect and its implementations, the second hearing device comprising: a second voice detector according to one of the first aspect and its implementations, configured to obtain one or more second microphone signals; and comprising a second noise suppressor; the first noise suppressor and the second noise suppressor are used in cooperation to perform the following operations: processing the one or more first microphone signals and the one or more second microphone signals to obtain a combined microphone signal; generating a modified combined microphone signal by selectively applying a gain to the voice signal in the combined microphone signal.
[0049] For example, the first hearing device can be used for one ear of the user, and the second hearing device can be used for the other ear of the user. In this case, the system according to the third aspect can ensure to provide the best hearing experience for the user.
[0050] In one embodiment, the first noise suppressor and the second noise suppressor are used to form a single noise suppressor.
[0051] Advantageously, the system can amplify or reduce ambient sound, but keep the user's own voice at (or close to) its original level through the noise suppressor for attenuating low-level noise, etc.
[0052] In one implementation of the system according to the third aspect, the first hearing device further comprises a first enhancer, and the second hearing device further comprises a second enhancer; the first enhancer and the second enhancer are used in cooperation to apply enhancement to the modified combined microphone signal.
[0053] In one embodiment, the first enhancer and the second enhancer are used to form a single enhancer.
[0054] In one implementation of the system according to the third aspect, the combined microphone signal is obtained by combining the one or more first microphone signals and the one or more second microphone signals; or obtained by beamforming; or obtained by selecting the one or more first microphone signals or the one or more second microphone signals (depending on which microphone signal has higher signal quality) as the combined microphone signal.
[0055] According to a fourth aspect, the present invention relates to a voice detection method, the method comprising: obtaining one or more microphone signals; obtaining a voice accelerometer (VAC) signal; identifying whether a fundamental tone exists in the VAC signal according to the one or more microphone signals; and if a fundamental tone is detected in the VAC signal, determining whether the fundamental tone is related to a voice signal.
[0056] The method according to the fourth aspect has the same advantages as the voice detector according to the first aspect, and can be extended by the corresponding implementation manners described above for the voice detector according to the first aspect.
[0057] According to a fifth aspect, the present invention relates to a computer program, comprising: program code for, when executed on a computer, performing the method according to the fourth aspect or any of its implementation manners.
[0058] According to a sixth aspect, the present invention relates to a non-transitory storage medium storing executable program code, which, when executed by a processor, causes the method according to the fourth aspect or any of its implementation manners to be performed.
[0059] It should be noted that all devices, elements, units and modules described in the present application can be implemented in software or hardware elements or any combination thereof. The steps performed by the various entities described in the present application and the functions to be performed by the various entities described are intended to mean that each entity is used to perform each step and function. Even in the description of the following specific embodiments, where the specific functions or steps to be performed by an external entity are not reflected in the description of the specific detailed elements of the entity performing the specific step or function, those skilled in the art should understand that these methods and functions can be implemented in the corresponding software or hardware elements, or in any combination of such elements. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In combination with the accompanying drawings, the following description of specific embodiments elaborates on the aspects and implementation manners of the present invention described above.
[0061] Figure 1 A schematic diagram of a voice detector provided by an embodiment of the present invention is shown;
[0062] Figure 2 The fundamental tone of a male voice provided by an embodiment of the present invention is shown;
[0063] Figure 3 The fundamental tone of a female voice provided by an embodiment of the present invention is shown;
[0064] Figure 4 A schematic diagram of a hearing device including a voice detector provided by an embodiment of the present invention is shown;
[0065] Figure 5 Shows signals processed by a hearing device including a voice detector provided by an embodiment of the present invention;
[0066] Figure 6 Shows a schematic diagram of a hearing device including a voice detector provided by an embodiment of the present invention;
[0067] Figure 7 Shows a schematic diagram of a system including a voice detector for a hearing device provided by an embodiment of the present invention;
[0068] Figure 8 Shows a schematic diagram of a voice detection method provided by an embodiment of the present invention. Detailed implementation
[0069] Figure 1 Shows a schematic diagram of a voice detector 100 provided by an embodiment of the present invention. The voice detector 100 is used for a hearing device 400 (see Figure 4 ), specifically, it can be a part of the hearing device 400. In some embodiments, the voice detector 100 can be a supplementary device connected to the hearing device 400.
[0070] The voice detector 100 is used to obtain one or more microphone signals 201a from one or more microphones 201. The one or more microphones 201 can be a part of the hearing device 400. Specifically, each microphone 201 can provide a microphone signal 201a to the voice detector 100.
[0071] In addition, the voice detector 100 is used to obtain a VAC signal 202a from a VAC 202. The VAC 202 can be a part of the hearing device 400.
[0072] In addition, the voice detector 100 is used to identify whether there is a fundamental tone in the VAC signal 202a according to the one or more microphone signals (represented by unit 101). In other words, the voice detector 100 can obtain information available for the fundamental tone detection from the one or more microphone signals 201a (in unit 102). For example, as described below, according to the signal power of the one or more microphone signals 201a, the fundamental tone in the VAC signal 202a can be detected with different sensitivities.
[0073] Specifically, the voice detector 100 can be used to: determine a first VAC threshold according to the one or more microphone signals 201a (in unit 101), for example, according to the signal power of the one or more microphone signals 201a; identify whether there is a pitch in the VAC signal 202a according to the first VAC threshold; wherein, for different detected signal powers of the one or more microphone signals 201a, the first VAC threshold can be different. As Figure 1 shown, specifically, the voice detector 100 can be used to: determine the first VAC threshold by comparing the signal power of the current frame of the one or more microphone signals 201a with the average signal power of multiple frames of the one or more microphone signals 201a (in unit 101, receiving the one or more microphone signals 201a as the input of the one or more microphones 201).
[0074] If a pitch is identified in the VAC signal 202a, the voice detector 100 is further used to determine whether the pitch is related to a voice signal. Specifically, the voice detector 100 can be used to determine whether the pitch is related to the voice signal according to a second VAC threshold (in unit 103). For example, if the signal power of the VAC signal 202a is higher than the second VAC threshold, the voice detector 100 can only determine that the pitch is related to the voice signal; additionally, if the frequency of the pitch is within a predefined frequency range, it is determined that the pitch is related to the voice signal.
[0075] Therefore, the VAC signal 202a that is typically generated inside the user's ear and the one or more microphone signals 201a that are typically generated outside the user's ear can be used to detect the user's own voice. A typical VAC 202 can be used to pick up vowels in the user's own voice, but can also pick up other sounds caused by movement (e.g., chewing). In addition, a typical implementation of the VAC 202 (e.g., a vision processing unit (VPU)) may be very sensitive to interference.
[0076] Therefore, the voice detector 100 can advantageously use the pitch in the VAC signal 202a, the signal power of the VAC signal 202a, and the signal power of the one or more microphone signals 201a to more reliably detect the user's own voice. Therefore, the embodiments of the present invention can very precisely detect the user's own voice and further distinguish it from other sounds in the user's mouth.
[0077] As described above, in addition to the fundamental tone, the corresponding signal power can also play a role in the self-voice detection performed by the voice detector 100. For example, the first VAC threshold can be modified according to the presence of the microphone signal. In fact, this situation can be monitored by calculating the average signal power of multiple frames of the one or more microphone signals 201a (in unit 101) and comparing the signal power of the current frame of the one or more microphone signals 201 with the average signal power. If there is a signal in the one or more microphone signals 201 (i.e., the current signal power is higher than the average signal power), the first VAC threshold for pitch detection can be lower; if there is no signal in the one or more microphone signals 201 (i.e., the current signal power is not higher than the average signal power), the first VAC threshold can be higher. This is because one's own voice should only be attenuated in the further processing when one's own voice can be heard in the one or more microphones 201. In other words, this advantageously reduces the possibility of false voice detection.
[0078] Secondly, if the signal power of the VAC signal 202a is low, there may be no own voice. This situation can be monitored in a similar way to the presence of a voice signal in the one or more microphone signals 201. It is worth noting that if there is a voice signal in the one or more microphone 201 but no voice signal in the VAC 202, the one or more microphone signals 201a can be the target signals that should be enhanced. In addition, there may be periodic interference in the VAC 202. However, they can be excluded because their power is constant, although the pitch detection may mark them as voice. Specifically, if the VAC signal 202a (input to unit 103) is higher than the second VAC threshold, the result of the own voice detection (OVD) is positive (O.V.D. = 1), otherwise it is not positive (O.V.D. = 0).
[0079] In addition, the voice detector 100 can be used to calculate the pitch according to the cepstral coefficients that reveal periodicity and harmonics by the following formula c Calculate the pitch:
[0080] c = IFFT(abs(FFT( x )) 2 )
[0081] where x represents a 30 ms frame of the voice signal (i.e., the VAC signal), and IFFT and FFT respectively represent the inverse fast Fourier transform and the fast Fourier transform of the signal x .
[0082] The fundamental pitch can also be calculated at a lower sampling rate (e.g., 2 kHz) only when the signal power of the VAC signal is high enough (i.e., higher than the first VAC threshold), so as to have medium complexity.
[0083] Figure 2 and Figure 3 respectively show the calculation of the cepstral coefficients of male speech and female speech. Finally, if the maximum cepstral coefficient is within the frequency range of [65 Hz, 320 Hz], the maximum cepstral coefficient is divided by the signal power c(0) (the first element of the vector corresponding to log(0) or infinite fundamental pitch), and compared with the first VAC threshold. A very low fundamental pitch (i.e., a frequency lower than 65 Hz) is considered to be chewing or other non-speech activities inside the user's mouth. c In one embodiment, since the audible device is for personal use, and the fundamental pitch of male users is relatively low and that of female speakers is relatively high, two counters are provided in the voice detector 100 to count the percentage of frames of the VAC signal 202a in which a clear fundamental pitch exists, so as to determine whether the fundamental pitch is a male fundamental pitch or a female fundamental pitch. After making the determination, only the female fundamental pitch or the male fundamental pitch will be further searched for. In other words, if it is determined that the fundamental pitch is a female fundamental pitch, for example, it can be further searched for only within the frequency range of [120 Hz, 320 Hz]. Alternatively, if it is determined that the fundamental pitch is a male fundamental pitch, for example, it can be further searched for only within the frequency range of [65 Hz, 160 Hz]. When speaking, the fundamental pitch is usually not constant but varies. In tonal languages such as Chinese, words have different meanings depending on the variation of the fundamental pitch during vowels, and for example, in English, the fundamental pitch rises in case of problems. For example, in Finnish, the voice fundamental pitch usually drops monotonically.
[0084] FIG. shows a schematic diagram of a hearing device 400 provided by an embodiment, including a voice detector 100 such as shown in FIG.
[0085] Figure 4 FIG. shows an embodiment including such as Figure 1 as shown in FIG.
[0086] Accordingly, the hearing device 400 includes the voice detector 100 and also includes a noise suppressor 401 configured to generate one or more modified microphone signals 401a by selectively applying a gain to the voice signals in the one or more microphone signals 201. The voice signals correspond to the voice signals in the VAC signal 202a detected by the voice detector 100. In other words, the noise suppressor 401 selectively applies the gain to the voice signals only when the voice detector 100 detects the voice signals in the VAC signal 202a and thus implicitly detects the voice signals in the one or more microphone signals 201.
[0087] The noise suppressor 401 may also be configured to generate the one or more modified microphone signals 401a by suppressing background noise signals in the one or more microphone signals 201a.
[0088] The hearing device 400 may further include an enhancer 402 configured to apply an enhancement to the one or more modified microphone signals, in particular an enhancement determined according to a user input. In other words, through the enhancer 402, the user can control the total signal power output by the hearing device 400, that is, the user can adjust the loudness.
[0089] Advantageously, the hearing device 400 can thus amplify or reduce ambient sounds, but keep the user's own voice at the original level by means of the noise suppressor 401 for attenuating low-level background noise.
[0090] In fact, in a natural environment, there are always some low-level background noises (distant traffic, air conditioners, machines, ovens, refrigerators, computers, etc.). The user generally does not notice these noises. However, when such noises are reproduced during natural listening, this situation is similar but not identical to the situation where the user is not wearing any audible device at all. Therefore, it is considered disturbing, and the noise suppressor 401 can be used to attenuate this noise.
[0091] In one embodiment, the user manually operates the enhancer 402 through a user interface (UI) 203. The user interface 203 can be the user interface of the hearing device 400 or can be connected or communicate with the hearing device 400. Similar to how the user can adjust the playback volume of the wearable device, the user can also adjust the level of ambient noise. At the same time, the hearing device 400 can be used to keep the user's own voice at the original level and push low-level background noise to a predefined level where it is barely audible. This can be achieved by modifying the noise suppressor 401. In one embodiment, the user manually operates the noise suppressor 401 through the user interface (UI) 203.
[0092] Specifically, the required enhancement can be achieved from the user interface (UI) 203. For example, if the enhancement is zero (0 dB), the signal power does not change, and the user's own voice does not require control or enhancement. If the enhancement is negative, the signal is attenuated; if the enhancement is positive, the signal power increases.
[0093] For example, this means that if no enhancement is applied, the gain is zero; if negative enhancement is applied, the gain is positive; if positive enhancement is applied, the gain is negative. After enhancement, the enhanced and modified microphone signal 402a is provided as an input to the speaker 204 for reproduction to the user. The speaker can be part of the hearing device 400 or connected to the hearing device (in which case the hearing device 400 can be an auxiliary device connectable to any type of speaker 204).
[0094] Self-voice control can be efficiently achieved through the voice detector 100. Whenever the enhancement is changed, the noise suppressor 401 is retuned and the relevant gain parameters are reinitialized. This has the advantage of being computationally efficient because most of the calculations are performed during the initialization process.
[0095] The voice signal can be the user's own voice, which is usually higher than any other signal and will become too loud after amplification. The user can also adapt the volume of their own voice to the ambient voice volume and the way the user hears their own voice.
[0096] To better illustrate how the hearing device 400 processes different sound signals, in Figure 5 the curve 501 represents the original signal detected by the voice detector 100, while the curves 502 and 503 represent the signals enhanced by ±6 dB by the enhancer 402. In Figure 5In it, the low-level noise is not enhanced (from 20 to 32 seconds), and the user's own voice is not enhanced either (from 38 to 40 seconds). It is worth noting that Figure 5 shows the curve values (y-axis in dB) varying with time (x-axis in seconds).
[0097] Generally speaking, it has the following advantages, that is, the user can naturally hear his own voice at the original level, and the low-level noise will not be amplified, although everything else will attenuate or become louder.
[0098] Figure 6 shows a schematic diagram of the hearing device 400 provided by the embodiment, especially the noise suppressor 401 and the enhancer 402.
[0099] For example, in this embodiment of the hearing device 400, the noise suppressor 401 processes the microphone signal in a 10 ms frame x . The noise suppressor 401 can be used to transfer the product of each frame and the gain G(t,ω) to the frequency domain through FFT, and transfer it back to the frequency domain through the inverse fast Fourier transform. The gain G(t,ω) depends on the power spectral density (PSD) P(t,ω) calculated by the noise suppressor 401 according to the FFT of the frame, the noise N(t,ω) at time t, and the frequency ω.
[0100] Ideally, for pure speech, G(t,ω)=1; for pure noise, G(t,ω)=0; for noisy speech, G(t,ω) is between the two, depending on the estimated speech and noise levels. In fact, the gain is limited below this value; in one embodiment, the gain is 0.25. In the case of using decibels as the unit, this corresponds to 0 dB attenuation for speech and -12 dB attenuation for pure noise. In this case, the maximum attenuation parameter or the noise suppressor parameter m is 12 dB.
[0101] If the speech signal is enhanced, the parameter m of the noise suppressor 401 can be modified. In the case of positive enhancement of x dB, the gain is limited below -(x + m) dB; in the case of negative enhancement of -x dB (x < m), the gain is limited below (x - m) dB, otherwise, the hearing device 400 can be used to turn off the noise reduction function.
[0102] Finally, in one embodiment, when the user's own voice control marks the user's voice activity, the hearing device 400 can be used to further modify the parameter of the noise suppressor 401. In the case of positive enhancement of x dB, the gain is limited to G ovd(t, ω) = min(G(t, ω), -x) above, so that the enhancer 402 can be used to enhance the noise suppression signal back to the original level, where the noise is suppressed by m dB. In the case of negative enhancement, the hearing device 400 can be used to modify the parameters of the noise suppressor 401 such that: for speech, G ovd (t, ω) = 10 x / 20 ; for noise, G ovd (t, ω) = 10 max (x-m,0) / 20 .
[0103] Then, the signal power of the one or more microphone signals 201a from the one or more microphones 201 can be adjusted according to the gain parameter (in the enhancer 402), and the enhanced and modified microphone signal 402a can be provided to the speaker 204 as an input.
[0104] Figure 7 A schematic diagram of a system 700 provided by an embodiment of the present invention is shown. The system 700 includes a first hearing device 400a (e.g., for one ear of a user) and a second hearing device 400b (e.g., for the other ear of the user), wherein the first hearing device 400a and the second hearing device 400b are shown as being mechanically connected. However, the hearing devices 400a and 400b can also be separate from each other.
[0105] The first hearing device 400a includes: a first voice detector 100a for obtaining one or more first microphone signals 201a; and includes a first noise suppressor 401b. The second hearing device 400b includes: a second voice detector 100b for obtaining one or more second microphone signals 201; and includes a second noise suppressor 401c. The functions of the voice detector 100a and the voice detector 100b can be the same. The functions of the voice detector 100a and the voice detector 100b can be related to the voice detector 100 as described above.
[0106] In addition, the first noise suppressor 401b and the second noise suppressor 401c are respectively used in cooperation to perform the following operations: processing the one or more first microphone signals 201a and the one or more second microphone signals 201a to obtain a combined microphone signal; generating a modified combined microphone signal 401a by selectively applying a gain to the voice signal in the combined microphone signal.
[0107] In one embodiment, the combined microphone signal is obtained by combining the one or more first microphone signals 201a and the one or more second microphone signals 201a; or by beamforming; or by selecting the one or more first microphone signals 201a or the one or more second microphone signals 201a (depending on which microphone signal has higher signal quality (or power)) as the combined microphone signal.
[0108] In one embodiment, the first noise suppressor 401b and the second noise suppressor 401c are used to form a single noise suppressor 401.
[0109] In yet another embodiment, the first hearing device 400a further includes a first enhancer 402b, and the second hearing device 400b further includes a second enhancer 402c; the first enhancer 402b and the second enhancer 402c are used in cooperation to apply enhancement to the modified combined microphone signal 401a.
[0110] In one embodiment, the first enhancer 402b and the second enhancer 402c are used to form a single enhancer 402.
[0111] Figure 8 A schematic diagram of the voice detection method 800 provided by the embodiment is shown. The method 800 can be executed by the voice detector 100 (see Figure 1 ), or can be executed by each of the voice detector 100a and the voice detector 100b ( Figure 7 ).
[0112] The method 800 includes the following steps: Step 801: Obtain one or more microphone signals 201a; Step 802: Obtain the VAC signal 202a; Step 803: Identify whether there is a pitch in the VAC signal 202a according to the one or more microphone signals 201a; Step 804: If a pitch is detected in the VAC signal, determine whether the pitch is related to a voice signal.
[0113] The present invention has been described in conjunction with various embodiments and implementation manners by way of example. However, those skilled in the art can understand and obtain other variations by practicing the present invention, studying the drawings, the present invention, and the appended claims. In the claims and the specification, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" does not exclude a plurality. A single element or other unit may fulfill the functions of several entities or items described in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not mean that the combination of these measures cannot be used effectively.
Claims
1. A voice detector (100) for a hearing device (400), characterized in that, The voice detector (100) is used for: Obtaining one or more microphone signals (201a); Obtaining a voice accelerometer VAC signal (202a); Identifying whether there is a pitch in the VAC signal (202a) according to the signal power of the one or more microphone signals (201a); If it is identified that there is a pitch in the VAC signal (202a), Then determining whether the pitch is related to the user's voice signal.
2. The voice detector (100) according to claim 1, characterized in that, It is also used for: Determining a first VAC threshold according to the one or more microphone signals (201a); Identifying whether there is the pitch in the VAC signal (202a) according to the first VAC threshold.
3. The voice detector (100) according to claim 2, characterized in that, It is also used for: Determining whether the pitch is related to the voice signal according to the determined second VAC threshold.
4. The voice detector (100) according to claim 3, characterized in that, For: Additionally, if the frequency of the pitch is within a predefined frequency range, determining that the pitch is related to the voice signal.
5. The voice detector (100) according to any one of claims 2 to 4, characterized in that The first VAC threshold is determined by comparing the signal power of the current frame of the one or more microphone signals with the average signal power of multiple frames of the one or more microphone signals.
6. The voice detector (100) according to claim 5, characterized in that If the signal power of the current frame of the one or more microphone signals is higher than the average signal power of the one or more microphone signals, the first VAC threshold has a higher value; and / or If the signal power of the current frame of the one or more microphone signals is equal to or lower than the average signal power of the one or more microphone signals, the first VAC threshold has a lower value.
7. The voice detector (100) according to any one of claims 1-4, 6, characterized in that, In order to identify whether there is the pitch in the VAC signal (202a), the voice detector (100) is used for: Determining whether the pitch detected in at least one frame of the VAC signal (202a) is a male pitch or a female pitch.
8. The voice detector (100) according to claim 7, wherein, In order to identify whether there is the pitch in the VAC signal (202a), the voice detector (100) is also used for: Searching for the determined male pitch in other frames of the VAC signal (202a); or Searching for the determined female pitch in other frames of the VAC signal (202a).
9. The voice detector (100) according to claim 3 or 4, characterized in that, In order to identify whether there is the pitch in the VAC signal, the voice detector (100) is also used for: Calculating one or more cepstral coefficients according to the VAC signal (202a); Calculating the pitch according to the one or more cepstral coefficients based on the first VAC threshold and the second VAC threshold.
10. The voice detector (100) according to claim 9, characterized in that The one or more cepstral coefficients are calculated according to the following formula: c = IFFT(abs(FFT(x))^2), where x represents the VAC signal, FFT represents the fast Fourier transform, and IFFT represents the inverse fast Fourier transform.
11. The voice detector (100) according to claim 9, characterized in that, To identify whether the fundamental tone exists in the VAC signal (202a), the voice detector (100) is further configured to: Determine the maximum cepstral coefficient corresponding to a certain frequency range; If the value obtained by dividing the maximum cepstral coefficient by the signal power is greater than the first VAC threshold, it is identified that the fundamental tone exists in the VAC signal (202a).
12. The voice detector (100) according to claim 10, characterized in that, To identify whether the fundamental tone exists in the VAC signal (202a), the voice detector (100) is further configured to: Determine the maximum cepstral coefficient corresponding to a certain frequency range; If the value obtained by dividing the maximum cepstral coefficient by the signal power is greater than the first VAC threshold, it is identified that the fundamental tone exists in the VAC signal (202a).
13. The voice detector (100) according to claim 11 or 12, characterized in that, For: If the normalized maximum cepstral coefficient of the current frame of the VAC signal (202a) is higher than the second VAC threshold, it is determined that the fundamental tone is related to the voice signal in the current frame of the VAC signal (202a); Wherein, the second VAC threshold is determined by the average signal power of multiple frames of the VAC signal (202a).
14. A hearing device (400), characterized in that, Comprising: The voice detector (100) according to any one of claims 1 to 12; A noise suppressor (401) for generating one or more modified microphone signals (401a) by selectively applying a gain to the voice signal in one or more microphone signals (201a); Wherein, the voice signal corresponds to the voice signal in the VAC signal (202a) detected by the voice detector (100).
15. The hearing device (400) according to claim 14, wherein The noise suppressor (401) is further configured to generate the one or more modified microphone signals (401a) by suppressing the background noise signal in the one or more microphone signals (201a).
16. The hearing device (400) according to claim 14 or 15, characterized in that, Further comprising: An enhancer (402) for applying enhancement to the one or more modified microphone signals (401a).
17. The hearing device (400) according to claim 16, characterized in that, The enhancement applied to the one or more modified microphone signals (401a) is enhancement determined according to user input.
18. The hearing device (400) according to claim 16, wherein According to the enhancement applied to the one or more modified microphone signals (401a), the gain applied to the voice signal is determined.
19. The hearing device (400) according to claim 17, wherein According to the enhancement applied to the one or more modified microphone signals (401a), the gain applied to the voice signal is determined.
20. The hearing device (400) according to claim 16, characterized in that, For: Select the gain according to the enhancement such that the signal power of the voice signal in the one or more modified microphone signals (401a) is equal to the signal power of the voice signal in the one or more microphone signals (201a).
21. The hearing device (400) according to any one of claims 17-19, characterized in that, For: The gain is selected according to the enhancement such that the signal power of the speech signal in the one or more modified microphone signals (401a) is equal to the signal power of the speech signal in the one or more microphone signals (201a).
22. The hearing device (400) according to claim 16, characterized in that if no enhancement is applied, the gain is zero; if negative enhancement is applied, the gain is positive; if positive enhancement is applied, the gain is negative.
23. The hearing device (400) according to any one of claims 17 - 20, characterized in that if no enhancement is applied, the gain is zero; if negative enhancement is applied, the gain is positive; if positive enhancement is applied, the gain is negative.
24. The hearing device (400) according to any one of claims 14, 15, 17, 18, 19, 20, characterized in that, Further comprising: one or more microphones (201) for generating the one or more microphone signals (201a); and / or VAC (202) for generating the VAC signal (202a).
25. The hearing device (400) according to claim 16, characterized in that, Further comprising: one or more microphones (201) for generating the one or more microphone signals (201a); and / or VAC (202) for generating the VAC signal (202a).
26. A system (700), characterized in that, Comprising: A first hearing device (400a), the first hearing device (400a) being a hearing device according to any one of claims 14 to 25, the first hearing device (400a) comprising: a first speech detector (100a), the first speech detector (100a) being a speech detector according to any one of claims 1 to 13 for obtaining one or more first microphone signals; and the first hearing device (400a) comprising a first noise suppressor (401b); A second hearing device (400b), the second hearing device (400b) being a second hearing device (400b) according to any one of claims 14 to 25, the second hearing device (400b) comprising: a second speech detector (100b), the second speech detector (100b) being a speech detector according to any one of claims 1 to 13, the second speech detector (100b) for obtaining one or more second microphone signals; and the second hearing device (400b) comprising a second noise suppressor (401c); The first noise suppressor (401b) and the second noise suppressor (401c) are used in cooperation to perform the following operations: - Process the one or more first microphone signals and the one or more second microphone signals to obtain a combined microphone signal; - Generate a modified combined microphone signal (401a) by selectively applying a gain to the speech signal in the combined microphone signal.
27. The system (700) according to claim 26, characterized in that the first hearing device (400a) further comprises a first enhancer (402b), and the second hearing device (400b) further comprises a second enhancer (402c); The first enhancer (402b) and the second enhancer (402c) are used in combination to apply enhancement to the modified combined microphone signal.
28. The system (700) according to claim 26 or 27, characterized in that, The combined microphone signal is obtained by combining the one or more first microphone signals and the one or more second microphone signals; or by beamforming; or by selecting the one or more first microphone signals or the one or more second microphone signals (depending on which microphone signal has higher signal quality) as the combined microphone signal.
29. A voice detection method (800), characterized in that, The method (800) includes: Obtaining one or more microphone signals (201a); Obtaining a voice accelerometer VAC signal (202a); Identifying whether there is a pitch in the VAC signal (202a) according to the signal power of the one or more microphone signals (201a); If a pitch is detected in the VAC signal (202a), Determining whether the pitch is related to the user's voice signal.
30. A computer program, characterized in that, Including: Program code for performing the method (800) according to claim 29 when running on a computer.
Citation Information
Patent Citations
System and method of detecting a user's voice activity using an accelerometer
US20140093091A1
System and method for performing automatic gain control using an accelerometer in a headset
US20170263267A1