Audio system and signal processing method for voice activity detection of ear-worn playback devices

By using voice activity detector (VAD) in headphones to detect user voice, the problem of difficulty in distinguishing user voice from third-person voice in adaptive ANC process is solved, and efficient adaptive ANC process and low-power operation are achieved.

CN113994423BActive Publication Date: 2025-06-10에이엠에스오스람아게
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202080022922.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-18
Filing Date
2020-03-17
Publication Date
2025-06-10
Estimated Expiration
2040-03-17

AI Technical Summary

Technical Problem

In the process of adaptive ANC, it is difficult to effectively distinguish between user voice and third-person voice, resulting in errors and false zero problems in the adaptive process.

Method used

Voice Activity Detector (VAD) is used to detect user voice through the relationship between two microphones, using simple parameters and algorithms for rapid detection, and does not rely on detecting the surrounding sound period between voices.

Benefits of technology

It realizes effective detection of user voice, maximizes adaptive bandwidth, reduces errors and false zero problems, and can run on low-power devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113994423B_ABST
    Figure CN113994423B_ABST
Patent Text Reader

Abstract

An audio system for an ear-worn playback device (HP) includes: a speaker (SP), an error microphone (FB_MIC) that mainly senses the sound output from the speaker (SP), and a feed-forward microphone (FF_MIC) that mainly senses ambient sound. The audio system further includes a voice activity detector (VAD) configured to record a feed-forward signal (FF) from the feed-forward microphone (FF_MIC). In addition, an error signal (ERR) from the error microphone (FB_MIC) is recorded. Detection parameters are determined based on the feed-forward signal (FF) and the error signal (ERR). The detection parameters are monitored and a voice activity state is set according to the detection parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an audio system and a signal processing method for voice activity detection in an ear-worn playback device, such as a headphone including a speaker, an error microphone, and a feedforward microphone. Background Art

[0002] Today, more and more headphones or earbuds are equipped with noise cancellation technology. For example, such noise cancellation technology is known as active noise cancellation or ambient noise cancellation, both abbreviated as ANC. ANC typically processes the recorded ambient noise to generate an anti-noise signal, which is then combined with the useful audio signal and played through the speaker of the headphone. ANC can also be used in other audio devices, such as a mobile phone or cellular phone. Various ANC methods use a feedback FB microphone or error microphone, a feedforward FF microphone, or a combination of a feedback microphone and a feedforward microphone. FF and FB ANC are implemented by tuning filters based on the given acoustics of the audio system.

[0003] Hybrid noise cancellation headphones are well known. For example, a microphone is placed in a space directly coupled to the eardrum, typically near the front of the headphone driver. This is referred to as a feedback FB microphone or error microphone. A second microphone, namely a feedforward FF microphone, is placed outside the headphone such that the second microphone is acoustically decoupled from the headphone driver.

[0004] Conventional ambient noise cancellation headphones are characterized in that the driver has air spaces in front of and behind it. The front space is partially constituted by the ear canal space of the user wearing the headphone. The front space typically consists of a vent covered with an acoustic resistor. The rear space is also typically characterized by a vent with an acoustic resistor. Usually, the front space vent acoustically couples the front space with the rear space. There are two microphones for each left and right channel. The error microphone or feedback FB microphone is placed near the driver such that the error microphone or feedback FB microphone detects the sound from the driver and the sound from the surrounding environment. The feedforward FF microphone is placed facing outwards from the rear of the unit such that the feedforward FF microphone detects the surrounding environment sound and a negligible sound from the driver.

[0005] With this arrangement, two forms of noise cancellation can be performed, namely feedforward FF and feedback FB. Both systems involve a filter placed between the microphone and the driver. The feedforward system detects the noise outside the headphone, processes the noise via the filter, and outputs an anti-noise signal from the driver such that the anti-noise signal and the noise signal overlap at the ear to produce noise cancellation. The signal path is as follows:

[0006] ERR = AE - AM.F.DE

[0007] Where ERR is the residual noise at the ear, AE is the acoustic transfer function from the environment to the ear, AM is the acoustic transfer function from the environment to the FF microphone, F is the FF filter, and DE is the acoustic transfer function from the driver to the ear. All signals are complex in the frequency domain and thus contain amplitude and phase components. Therefore, it can be concluded that for perfect noise cancellation, ERR tends to zero:

[0008]

[0009] However, in practice, the acoustic transfer functions may change depending on the wearing situation of the headphones. For leaking earbuds, there may be highly variable leakage that acoustically couples the front space to the surrounding environment, and the transfer functions AE and DE will essentially change, such that the FF filter needs to be adapted according to the acoustic signal in the ear canal to minimize the error. Unfortunately, when the headphone user speaks, the signal at the microphone becomes mixed with the bone-conducted voice signal and may cause errors and false zeros in the adaptation process. Summary of the Invention

[0010] The aim is to provide an audio system and a signal processing method for voice activity detection that allow improved voice activity detection, for example, detecting the voice present in the ear canal of the user of the audio system.

[0011] These aims are achieved by the subject matter of the independent claims. Further improvements and embodiments are described in the dependent claims.

[0012] It should be understood that any feature associated with any one embodiment can be used alone, or in combination with other features described herein, or in combination with one or more features of any other embodiment, or in any combination with any other embodiment, unless described as an alternative. Additionally, equivalent parts and modifications not described below can also be employed without departing from the scope of the audio system and method for voice activity detection defined by the appended claims.

[0013] The following relates to improved concepts in the field of environmental noise cancellation. The improved concepts allow, for example, the implementation of voice activity detection in a playback device, such as headphones that require a first-person voice activity detector, which may be necessary for an adaptive ANC process, acoustic switch ear detection, and voice commands. The improved concepts can be applied to the adaptive ANC of leaking earbuds. The term "adaptive" refers to adapting the anti-noise signal according to the leakage that acoustically couples the front space of the device to the surrounding environment. The voice activity detector uses the relationship between two microphones to detect the user's voice rather than a third-person voice and uses the relationship between two microphones to detect the user's voice in the headphone scenario. The improved concepts also focus on simple parameters to keep the processing volume to a minimum.

[0014] The improved concept can avoid detecting third-person voice, which means that in the context of adaptive ANC headphones, adaptation only occurs when the user is speaking (i.e., the first person), rather than a third party, thus maximizing the adaptive bandwidth. It can only detect bone-conducted voice.

[0015] The improved concept can be implemented with a simple algorithm, which essentially means that it can operate at a lower power (on devices with lower specifications) than some algorithms.

[0016] The improved concept does not rely on detecting the ambient sound period between voices as a reference (e.g., the coherence method). Essentially, its reference is essentially the known phase relationship between microphones. Therefore, it can quickly determine whether there is sound.

[0017] In at least one embodiment, an audio system for an ear-worn playback device includes a speaker, an error microphone that mainly senses the sound output from the speaker, and a feedforward microphone that mainly senses ambient sound. The audio system further includes a voice activity detector configured to perform the following steps, including recording a feedforward signal from the feedforward microphone and recording an error signal from the error microphone. Determining a detection parameter based on the feedforward signal and the error signal. Monitoring the detection parameter and setting a voice activity state based on the detection parameter.

[0018] In at least one embodiment, the detection parameter is based on the ratio of the feedforward signal to the error signal.

[0019] In at least one embodiment, the detection parameter is further based on a sound signal.

[0020] In at least one embodiment, the detection parameter is the amplitude difference between the feedforward signal and the error signal. The detection parameter can indicate ANC performance. For example, ANC performance is determined by the ratio of the amplitudes between the microphones.

[0021] In at least one embodiment, the detection parameter is the phase difference between the error signal and the feedforward signal.

[0022] In at least one embodiment, the audio system further includes an adaptive noise cancellation controller coupled to the feedforward microphone and the error microphone. The adaptive noise cancellation controller is configured to perform noise cancellation processing based on the feedforward signal and / or the error signal. A filter is coupled to the feedforward microphone and the speaker and has a filter transfer function determined by the noise cancellation processing.

[0023] In at least one embodiment, the noise cancellation processing includes feedforward noise cancellation processing or feedback noise cancellation processing or both feedforward noise cancellation processing and feedback noise cancellation processing.

[0024] In at least one embodiment, a detection parameter indicates the performance of a noise cancellation process.

[0025] In at least one embodiment, a voice activity detector process determines one of the following voice activity states: false, true, or likely. A detection state equal to "true" indicates that voice is detected. A detection state equal to "false" indicates that no voice is detected. A detection state equal to "likely" indicates that voice is probable.

[0026] In at least one embodiment, a voice activity detector controls an adaptive noise cancellation controller based on the voice activity state.

[0027] In at least one embodiment, controlling the adaptive noise cancellation controller includes terminating the adaptation of a noise cancellation signal of the noise cancellation process when the voice activity state is set to "true" and / or "likely". When the voice activity state is set to "false", the adaptation of the noise cancellation signal continues.

[0028] In at least one embodiment, in a first operation mode, a voice activity detector analyzes a phase difference between a feedforward signal and an error signal. The voice activity state is set based on the analyzed phase difference.

[0029] In at least one embodiment, when the detection parameter is greater than or exceeds a first threshold, the first operation mode is entered, that is, generally speaking, the difference between the detection parameter and the first threshold is considered. Hereinafter, the term "exceeds" is considered equivalent to "greater than" or "more than".

[0030] In at least one embodiment, the phase difference is monitored in the frequency domain. The phase difference is analyzed based on an expected transfer function to record a deviation from the expected transfer function at at least some frequencies. The voice activity state is set based on the recorded deviation.

[0031] In at least one embodiment, voice is detected by identifying peaks in the phase difference in the frequency domain.

[0032] In at least one embodiment, the analyzed phase difference is compared with an expected phase difference. When the analyzed phase difference is less than the expected phase difference, the voice activity state is set to "false", otherwise it is set to "true". That is, generally speaking, the difference between the analyzed phase difference and the expected phase difference is considered not and should not exceed a predetermined value or range of values.

[0033] In at least one embodiment, in a second operation mode, a voice activity detector analyzes a pitch level of an error signal and sets the voice activity state based on the analyzed pitch level.

[0034] In at least one embodiment, the second operation mode is entered when the detection parameter is less than the first threshold.

[0035] In at least one embodiment, the analyzed pitch level is compared with an expected pitch level. When the analyzed pitch level exceeds the expected pitch level, the voice activity state is set to "true", otherwise it is set to "false". That is to say, generally speaking, the difference between the analyzed pitch level and the expected pitch level is considered not to and should not exceed a predetermined value or a range of values.

[0036] In at least one embodiment, in a third operation mode, the voice activity detector monitors a parameter, expressed as a short-term parameter, in a first time period and monitors a parameter, expressed as a long-term parameter, in a second time period. The first time period is shorter in time than the second time period. In addition, the voice activity detector combines the short-term parameter and the long-term parameter to obtain a combined detection parameter and sets the voice activity state according to the combined detection parameter. In at least one embodiment, the third mode can operate independently of the first two modes.

[0037] In at least one embodiment, the short-term parameter and the long-term parameter correspond to energy levels. When the change in the relative energy level exceeds a second threshold, the voice activity state is set to "probable".

[0038] In at least one embodiment, in a fourth operation mode, the voice activity detector determines whether a desired sound signal is active. If no sound signal is active, the voice activity detector enters the first operation mode or the second operation mode. If the sound signal is active and if a first threshold exceeds the analyzed detection parameter, the voice activity detector enters the second mode of operation, or if the sound signal is active and if the analyzed detection parameter exceeds the first threshold, the voice activity detector enters a combined first operation mode and second operation mode. In other words, if there is music, the voice activity detector can enter the second operation mode or the combined operation mode based on the detection parameter (e.g., ANC approximation).

[0039] In at least one embodiment, in the combined first operation mode and second operation mode, the voice activity detector analyzes the pitch level of the error signal and analyzes the phase difference between the feedforward signal and the error signal. In addition, the voice activity detector sets the voice activity state according to the analyzed phase difference and the analyzed pitch level.

[0040] In at least one embodiment, in a combined first and second operating mode, the analyzed tone level is compared with an expected tone level, and the analyzed phase difference is compared with an expected phase difference. When the analyzed tone level exceeds the expected tone level, and further, when the analyzed phase difference exceeds the expected phase difference, the voice activity state is set to "true". When the expected tone level exceeds the analyzed tone level, and / or when the expected phase difference exceeds the analyzed phase difference, the voice activity state is set to "false".

[0041] In at least one embodiment, the audio system includes an ear-worn playback device.

[0042] In at least one embodiment, an adaptive noise cancellation controller, a voice activity detector, and / or a filter are included in the housing of the playback device.

[0043] In at least one embodiment, the playback device is a headphone or an earbud.

[0044] In at least one embodiment, the headphone or earbud is designed to have a predefined acoustic leakage between the body of the headphone or earbud and the user's head when worn.

[0045] In at least one embodiment, the playback device is a mobile phone.

[0046] In at least one embodiment, the adaptive noise cancellation controller, the voice activity detector, and / or the filter are integrated into a general-purpose device.

[0047] In at least one embodiment, if the playback device is worn in the user's ear, the device has a front space and a rear space on either side of the driver, wherein the front space at least partially includes the user's ear canal. An error microphone is arranged in the playback device such that the error microphone is acoustically coupled to the front space. A feedforward microphone is arranged in the playback device such that the feedforward microphone faces outward from the rear space.

[0048] In at least one embodiment, the playback device includes a front vent with or without a first acoustic resistor that couples the front space to the surrounding environment. Additionally or alternatively, the playback device includes a rear vent with or without a second acoustic resistor that couples the rear space to the surrounding environment.

[0049] In at least one embodiment, the playback device includes a vent that couples the front space to the rear space.

[0050] A signal processing method for voice activity detection can be applied to an ear-worn playback device, which includes a speaker, an error microphone that senses the sound output from the speaker and the ambient sound, and a feedforward microphone that mainly senses the ambient sound. The method can be performed by means of a voice activity detector. In at least one embodiment, the method includes the steps of recording a feedforward signal from the feedforward microphone and recording an error signal from the error microphone. Detection parameters are determined based on the feedforward signal and the error signal. The detection parameters are monitored, and the voice activity state is set according to the detection parameters.

[0051] Other embodiments of the method can be easily derived from various embodiments and implementations of the audio system, and vice versa.

[0052] In all the above embodiments, ANC can be performed using both digital and / or analog filters. All audio systems can also include feedback ANC. The processing and recording of various signals are preferably performed in the digital domain.

[0053] According to one aspect, a noise-canceling ear-worn device includes: a driver that has a space in front of and behind it such that the front space is at least partially constituted by the ear canal; and an error microphone that is acoustically coupled to the front space and detects ambient noise and the driver signal; a feedforward (FF) microphone that faces outward from the rear space and detects ambient noise and a negligible portion of only the driver signal, whereby the feedforward FF microphone is coupled to the driver via a filter, causing the driver to output a signal that at least partially cancels the noise at the error microphone; and further includes a processor that monitors the phase difference between the two microphones, and the phase difference triggers the voice activity stage state according to the condition of the phase difference.

[0054] According to another aspect, the device as described above monitors the phase difference in the time domain, as well as the deviation from the expected transfer function at some frequencies, rather than others indicating that speech has occurred.

[0055] According to another aspect, a time-domain process is run to mark possible speech detection situations, and the time-domain process can act faster than the frequency-domain process.

[0056] According to another aspect, a second process is run to detect tones in the ambient signal.

[0057] According to another aspect, the second process is run in the frequency domain.

[0058] According to one aspect, an audio system for an ear-worn playback device (HP) includes:

[0059] - a speaker (SP),

[0060] - An error microphone (FB_MIC) that senses the sound output from the speaker and the ambient sound (SP), and

[0061] - A feedforward microphone (FF_MIC) that mainly senses the ambient sound,

[0062] wherein the audio system includes a voice activity detector (VAD) configured to:

[0063] - Record the feedforward signal (FF) from the feedforward microphone (FF_MIC),

[0064] - Record the error signal (ERR) from the error microphone (FB_MIC),

[0065] - Determine at least one detection parameter based on the feedforward signal (FF) and the error signal (ERR), and

[0066] - Monitor at least one detection parameter and set the voice activity state according to at least one detection parameter.

[0067] According to one aspect, the detection parameter is based on the ratio of the feedforward signal to the error signal.

[0068] According to one aspect, the detection parameter is the phase difference between the error signal and the feedforward signal.

[0069] According to one aspect, the detection parameter is also based on the sound signal (MUS).

[0070] According to one aspect, the voice activity detector (VAD) is configured to remove the sound signal (MUS) from the error signal (ERR).

[0071] According to one aspect, the detection parameter is the phase difference between the feedforward signal (FF) and the error signal (ERR).

[0072] According to one aspect, the audio system further includes:

[0073] - An adaptive noise cancellation controller (ANCC) coupled to the feedforward microphone (FF_MIC) and the error microphone (FB_MIC), the adaptive noise cancellation controller (ANCC) being configured to perform noise cancellation processing according to the feedforward signal (FF) and / or the error signal (ERR), and

[0074] - A filter (FL) coupled to the feedforward microphone (FF_MIC) and the speaker (SP), the filter having a filter transfer function (F) determined by the noise cancellation processing.

[0075] According to one aspect, the noise cancellation processing includes feedforward noise cancellation processing or feedback noise cancellation processing or both feedforward noise cancellation processing and feedback noise cancellation processing.

[0076] According to one aspect, the detection parameter indicates the performance of the noise cancellation processing.

[0077] According to one aspect:

[0078] - The voice activity detector process determines one of the following voice activity states: false, true, or likely,

[0079] - A voice activity state equal to true indicates that voice is detected, and

[0080] - A voice activity state equal to false indicates that voice may be detected.

[0081] According to one aspect, a voice activity detector (VAD) controls an adaptive noise cancellation controller (ANCC) according to the voice activity state.

[0082] According to one aspect, the control of the adaptive noise cancellation controller (ANCC) includes:

[0083] - Terminating the adaptation of the noise cancellation signal when the voice activity state is set to true and / or likely, and

[0084] - Continuing the adaptation of the noise cancellation signal when the voice activity state is set to false.

[0085] According to one aspect, in a first operation mode, the voice activity detector (VAD):

[0086] - Analyzes the phase difference between the feedforward signal (FF) and the error signal (ERR), and

[0087] - Sets the voice activity state according to the analyzed phase difference.

[0088] According to one aspect, when the detection parameter is greater than a first threshold, enter the first operation mode.

[0089] According to one aspect:

[0090] - Monitor the phase difference in the frequency domain,

[0091] - Analyze the phase difference based on the expected transfer function to record the deviation from the expected transfer function at at least some frequencies, and

[0092] - Set the voice activity state according to the recorded deviation.

[0093] According to one aspect, voice is detected by identifying peaks in the frequency domain phase response.

[0094] According to one aspect:

[0095] - Compare the analyzed phase difference with the expected phase difference, and

[0096] - When the analyzed phase difference is less than the expected phase difference, set the voice activity state to false, otherwise set it to true.

[0097] According to one aspect, in the second operating mode, a voice activity detector (VAD):

[0098] - Analyze the pitch level of the error signal (ERR), and

[0099] - Set the voice activity state according to the analyzed pitch level.

[0100] According to one aspect, enter the second operating mode when the first threshold is less than the detection parameter.

[0101] According to one aspect:

[0102] - Compare the analyzed pitch level with the expected pitch level,

[0103] - When the analyzed pitch level exceeds the expected pitch level, set the voice activity state to true, otherwise set it to false.

[0104] According to one aspect, in the third operating mode, a voice activity detector (VAD):

[0105] - Monitor the detection parameter within a first time period, represented as a short-term parameter, and monitor the detection parameter within a second time period, represented as a long-term parameter, wherein the first time period is shorter in time than the second time period,

[0106] - Combine the short-term parameter and the long-term parameter to obtain a combined detection parameter, and

[0107] - Set the voice activity state according to the combined detection parameter.

[0108] According to one aspect, in the third operating mode:

[0109] - The short-term parameter and the long-term parameter are equivalent to energy levels, and

[0110] - When the change in the relative energy level exceeds a second threshold, set the voice activity state to probable.

[0111] According to one aspect, in the fourth operating mode, a voice activity detector (VAD):

[0112] - Determine whether the sound signal (MUS) is active,

[0113] - If no voice signal (MUS) is active, enter the first operating mode or the second operating mode.

[0114] - If the voice signal (MUS) is active, and if the detection parameter is less than the first threshold, enter the second mode of operation, or

[0115] - If the voice signal (MUS) is active, and if the analyzed phase difference exceeds the first threshold, enter the combined first and second operating modes.

[0116] According to one aspect, in the combined first and second operating modes, a voice activity detector (VAD):

[0117] - Analyzes the pitch level of the error signal (ERR) and analyzes the phase difference between the feed-forward signal (FF) and the error signal (ERR), and

[0118] - Sets the voice activity state according to the analyzed phase difference and the analyzed pitch level.

[0119] According to one aspect, in the combined first and second operating modes:

[0120] - Compares the analyzed pitch level with the expected pitch level and compares the analyzed phase difference with the expected phase difference,

[0121] - When the analyzed pitch level is less than the expected pitch level and the analyzed phase difference is less than the expected phase difference, sets the voice activity state to false, and

[0122] - When the analyzed pitch level exceeds the expected pitch level and the analyzed phase difference exceeds the expected phase difference, sets the voice activity state to true.

[0123] According to one aspect, the audio system includes an ear-worn playback device.

[0124] According to one aspect, an adaptive noise cancellation controller (ANCC), a voice activity detector (VAD), and / or a filter (FL) are included in the housing of the playback device.

[0125] According to one aspect, the playback device is a headphone or an earplug.

[0126] According to one aspect, the headphone or earplug is designed to have a predefined acoustic leakage between the body of the headphone or earplug and the user's head when worn.

[0127] According to one aspect, the playback device is a mobile phone. According to one aspect, an Adaptive Noise Canceling Controller (ANCC), a Voice Activity Detector (VAD), and / or a Filter (FL) are integrated into a common driver (DRV).

[0128] According to one aspect, if the playback device is worn in the user's ear,

[0129] - the device has a front space and a rear space, where the front space at least partially includes the user's ear canal,

[0130] - an error microphone is provided in the playback device such that the error microphone is acoustically coupled to the front space, and

[0131] - a feedforward (FF) microphone is provided in the playback device such that the feedforward (FF) microphone faces outward from the rear space.

[0132] According to one aspect, the playback device includes:

[0133] - a front vent with or without a first acoustic resistor that couples the front space to the surrounding environment, and / or

[0134] - a rear vent with or without a second acoustic resistor that couples the rear space to the surrounding environment.

[0135] According to one aspect, the playback device includes a vent that couples the front space to the rear space.

[0136] According to one aspect, a signal processing method for voice activity detection in an ear-worn playback device (HP), the ear-worn playback device including a speaker (SP), an error microphone (FB_MIC) that mainly senses the sound output from the speaker (SP), and a feedforward microphone (FF_MIC) that mainly senses the surrounding ambient sound, the method includes the following steps:

[0137] - recording a feedforward signal (FF) from the feedforward microphone (FF_MIC),

[0138] - recording an error signal (ERR) from the error microphone (FB_MIC),

[0139] - determining a detection parameter based on the feedforward signal (FF) and the error signal (ERR), and

[0140] - monitoring the detection parameter and setting a voice activity state based on the detection parameter. Description of the Drawings

[0141] The improved concept is described in more detail below with reference to the drawings. In all the drawings, elements having the same or similar functions are denoted by the same reference numerals. Therefore, their descriptions will not be repeated in the following drawings.

[0142] In the drawings:

[0143] Figure 1 A schematic diagram of a headset is shown,

[0144] Figure 2 A block diagram of a general adaptive ANC system is shown,

[0145] Figure 3 An example representation of a "leaky" mobile phone is shown,

[0146] Figure 4 An example representation of a "leaky" earplug is shown,

[0147] Figure 5 The ERR(AE) and FF(AM) signal paths with respect to ambient noise are shown,

[0148] Figure 6 The ERR(BE) and FF(BM) signal paths of bone-conducted voice sounds are shown,

[0149] Figure 7 The frequency-phase response of the ERR / FF transfer function is shown,

[0150] Figure 8A and Figure 8B An ANC performance graph is shown,

[0151] Figure 9 An operating mode for rapid voice detection is shown, and

[0152] Figure 10 A flowchart of possible operating modes of a voice activity detector is shown. Detailed Description

[0153] Figure 1A schematic illustration of an ANC-enabled playback device in the form of headphones HP is shown, which in this example is designed as an ear-hook headphone or an ear-cup headphone. Only a part of the headphones HP corresponding to a single audio channel is shown. However, it will be obvious to those skilled in the art to extend it to stereo headphones. The headphones HP include a housing HS that houses a speaker SP, a feedback noise microphone or error microphone FB_MIC, and an ambient noise microphone or feed-forward microphone FF_MIC. The error microphone FB_MIC is particularly oriented or arranged such that it records the ambient noise and the sound played through the speaker SP. Preferably, the error microphone FB_MIC is arranged at a position close to the speaker, such as near the edge of the speaker SP or near the diaphragm of the speaker. The ambient noise / feed-forward microphone FF_MIC is particularly oriented or arranged such that it mainly records the ambient noise from outside the headphones HP. The error microphone FB_MIC can be used according to an improved concept to provide an error signal for voice activity detection.

[0154] In Figure 1 an embodiment, a sound control processor SCP including an adaptive noise cancellation controller ANCC is located within the headphones HP for performing various signal processing operations, examples of which will be described in the following disclosure. The sound control processor SCP can also be placed outside the headphones HP or within the headphones HP cable, for example, in an external device located in a mobile handheld device or a mobile phone.

[0155] Figure 2 A block diagram of a general adaptive ANC system is shown. The system includes an error microphone FB_MIC and a feed-forward microphone FF_MIC, both of which provide their output signals to the adaptive noise cancellation controller ANCC of the sound control processor SCP. The noise signal recorded by the feed-forward microphone FF_MIC is also provided to a feed-forward filter for generating an anti-noise signal output via the speaker SP. At the error microphone FB_MIC, the sound output from the speaker SP is combined with the ambient noise and recorded as an error signal ERR that includes the remaining part of the ambient noise after ANC. The adaptive noise cancellation controller ANCC uses this error signal ERR to adjust the filter response of the feed-forward filter. A voice activity detector VAD is coupled to the adaptive noise cancellation controller ANCC, the feed-forward microphone FF_MIC, and the error microphone FB_MIC.

[0156] For example, the embodiments are characterized in that the earplug EP has a driver, a front air space acoustically coupled to the front of the driver and partly constituted by the ear canal EC space, a rear space acoustically coupled to the back of the driver, a front vent that couples the front space to the ambient environment with or without an acoustic resistor, and a rear vent that couples the rear space to the ambient environment with or without an acoustic resistor. The front vent may be replaced by a vent that couples the front space and the rear space. When worn, the earplug EP may or may not have an acoustic leak between the front space and the ear canal space.

[0157] The error microphone FB_MIC may be placed such that it detects signals from the front of the driver and the ambient environment, and the feedforward FF microphone FF_MIC may be placed such that it detects ambient sound with a negligible part of the driver signal. When worn, the FF microphone is acoustically placed upstream of the error microphone FB_MIC with respect to ambient noise, and downstream of the error microphone with respect to bone conduction sound emitted from the ear canal wall.

[0158] The earplug EP may be characterized by FF, FB, or FF and FB noise cancellation. The noise cancellation at least partly adapts to changes in acoustic leakage. The bone conduction speech signal affects both microphone signals such that, in the presence of speech, the adaptation finds a suboptimal solution. Therefore, the adaptation must be stopped whenever the user speaks.

[0159] Both the FF microphone signal FF and the error microphone signal ERR are fed into a voice activity detector VAD, which analyzes these two signals to determine whether the user is speaking. The VAD returns three states: speech possible, speech false, and speech true. These states are passed to an adaptive noise cancellation controller ANCC, which makes a decision to stop adaptation, restart adaptation, or take no action.

[0160] The VAD runs three or four operating modes, such as two slow and one fast. The fast process detects a short-term increase in the level of the error microphone relative to the FF microphone. The fast process also detects a short-term increase in the FF microphone signal. If the short-term increase in the error microphone relative to the FF microphone exceeds a first threshold FT1, and the short-term increase in the FF microphone signal is lower than a second threshold FT2, then the VAD sets the state: speech possible. Then, the adaptive noise cancellation controller ANCC pauses adaptation in response.

[0161] One of the two slow processes runs according to an ANC performance approximation, which is the ratio of the long-term energy of the error microphone to the long-term energy of the FF microphone. If the ANC performance is greater than (worse than) the ANC threshold ANCT, which is a detection parameter, then the phase difference process that analyzes the phase difference between the two microphones or the first mode of operation is run. If the ANC performance is lower than (better than) ANCT, then the second mode of operation, which is the tone process that analyzes the tone of the error microphone, is run. The phase difference process and the tone process return a single metric, and this metric is tested against the threshold PDT for the phase difference or TONT for the tone. For example, the threshold is derived from the expected transfer function.

[0162] The phase difference process can perform a fast Fourier transform FFT on the error and FF microphone signals and calculate the phase difference between them. Before performing the FFT, the error signal and the FF signal can be downsampled to maximize the FFT resolution for a given processing volume.

[0163] The phase difference is calculated by dividing the two FFTs (ERR / FF) and taking the argument of the result. The smoothness of the resulting phase difference can be analyzed by various methods:

[0164] · Split the phase difference into several parts, calculate the local variance of each part, and then sum the results of each part to provide a single number for the variance.

[0165] · Apply linear regression to the data and calculate the squared deviation of each data point from the equivalent point in the resulting linear regression. Then sum the resulting deviations.

[0166] · Apply regression to an S-curve based on the Boltzmann equation and calculate the squared deviation of each data point from the resulting S-curve. Then sum the resulting deviations.

[0167] · Split the phase difference into several parts, calculate the local linear regression and calculate the squared deviation of each point from the equivalent linear regression point. Then sum the resulting deviations.

[0168] · High-pass filter the phase difference points to give a measure of smoothness, and then calculate the RMS energy or an equivalent measure of the resulting data.

[0169] · Calculate the rise and fall amplitudes of all the peaks, average each adjacent rise and fall amplitude to create a peak amplitude vector, consider small peaks below a cut-off value as noise and sum the remaining peaks.

[0170] The tone can be calculated in the frequency domain by taking the absolute value of the FFT of the error microphone FB_MIC signal ERR and calculating the measure of the peaks by using any of the metrics listed above for the phase difference variation.

[0171] The FFT for phase difference or pitch calculation can be replaced by several DFTs calculated at a predetermined frequency.

[0172] The phase difference or pitch can be calculated using any of the above methods, where the FFT is replaced by the energy level of the signal filtered by the Goertzel algorithm.

[0173] The phase difference can be calculated in the time domain by filtering and subtracting the signals from two microphones. If the phase difference exceeds a threshold, speech is assumed to be present.

[0174] The pitch can also be calculated in the time domain, for example by looking at zero crossings. Over a period of time, the linear regression of the zero-crossing-sample index can be calculated. If the squared deviation relative to the resulting regression is below a threshold, the signal is called pitched. If the deviation is above the threshold, the zero crossings are assumed to be random and the signal is not pitched. The input signal to the algorithm can be filtered to avoid the possibility of detecting pitch at frequencies outside the speech band.

[0175] The phase difference or pitch metrics can be averaged, or the PDT and TONT can be replaced with upper and lower threshold values (PDT1, PDT2, TONT1, TONT2) to apply hysteresis for improved output.

[0176] If the resulting pitch level or phase difference smoothness is higher than a set threshold, the speech state is set to true. ANCC stops adapting. If either parameter is below the set threshold, the speech state is set to false. ANCC restarts adapting.

[0177] If the desired signal (i.e., music) is played via the driver, both the pitch level and phase difference smoothness metrics must be higher than their respective thresholds for the VAD to set the speech state to true when the ANC performance approximation is higher than ANCT. This reduces false alarms triggered by music.

[0178] If the earbuds are a pair of left and right headphones, only one VAD needs to be run on one ear to set the speech state to possible, false, or true for both ears. When one earbud is removed from one ear and for the case of the ear on which the VAD is running, the VAD will switch to the other ear. This is done, for example, by reading the state of the ear-off detection module.

[0179] When the ANC performance approximation approaches ANCT, the phase difference VAD metric will return more false positives compared to the case where the ANC performance approximation is much higher (or lower) than ANCT. This is due to the non-smooth phase difference caused by the filter becoming closer to the acoustics. In this case, the false positives will slow down the adaptation speed, but this is acceptable because the ANC performance is close to the optimal zero value. If an earbud is removed from the ear, the VAD switches to the other ear and then the earbud is reinserted, then for the ear that has just been reinserted, although its ANC performance may be poor, the adaptation may be slow. In this case, to optimize the adaptation, the VAD is set to the ear in the in-ear state with the worst ANC performance approximation.

[0180] To know the ANC performance approximations on both sides, the fast VAD process must run simultaneously on the left and right ears.

[0181] Those skilled in the art should understand that there are many processes that can be used to detect peaks and valleys in the frequency domain and time domain. The improved concept is not limited to those shown here.

[0182] Now referring Figure 3 , another example of an audio system with noise cancellation enabled is presented. In this example embodiment, the system is formed by a mobile device, such as a mobile phone MP, which includes a playback device having a speaker SP, a feedback or error microphone FB_MIC, an ambient noise or feedforward microphone FF_MIC, and an adaptive noise cancellation controller ANCC for performing in particular ANC and / or other signal processing during operation.

[0183] In another embodiment not shown, headphones HP (such as Figure 1 or Figure 5 shown) can be connected to the mobile phone MP, where the signals from the microphones FB_MIC, FF_MIC are transmitted from the headphones to the mobile phone MP (such as the processor PROC of the mobile phone) for generating an audio signal to be played through the speaker of the headphones. For example, depending on whether the headphones are connected to the mobile phone, ANC is performed using the internal components of the mobile phone (i.e., the speaker and microphone) or using the speaker and microphone of the headphones, thus using different sets of filter parameters in each case.

[0184] Figure 4 An example representation of a "leaky" earbud is shown, i.e., the earbud is characterized by some leakage between the ambient environment and the ear canal EC. In particular, there is a sound path between the ambient environment and the ear canal EC, which is represented as "acoustic leakage" in the figure.

[0185] The proposed concept analyzes the signals at the error microphone FB_MIC and the FF microphone FF_MIC to infer whether there is speech in the ear canal EC. FF noise cancellation can be processed as described in the overview section, such that the signal at the FF microphone FF is the ambient noise at the FF microphone:

[0186] FF = AM

[0187] The signal ERR at the error microphone can be expressed as:

[0188] ERR = AE - AM.F.DE

[0189] Dividing the two gives a set of responses:

[0190]

[0191] In the frequency domain, all signals are complex and thus contain amplitude and phase components. It can be seen that the ratio of the two microphone signals ERR and FF is driven in part by the ratio of the acoustic transfer functions AE and AM.

[0192] Generally speaking, humans hear their own voices through three pathways. The first is the air conduction path, where sound travels from the mouth to the ear and is heard in the same way as ambient noise. The second is via the bone conduction pathway, which stimulates the inner part of the ear without becoming air-conducted. The third is via the bone conduction pathway, which passes through the ear canal wall and into the air, stimulating the eardrum in the same way as the surrounding ambient sound. It is this third path that disrupts the error signal and causes problems with the adaptive noise cancellation parameters of the headphones.

[0193] A voice activity detector can be used to detect speech from the person wearing the headphones rather than speech from an ambient noise source (i.e., detect the user's speech but ignore speech signals from third parties). The transfer function of bone-conducted sound varies from person to person and also depends on the way the headphones are worn (e.g., due to the occlusion effect). Therefore, it may not be possible to continue adapting by using a general bone conduction transfer function while there is speech. Therefore, when there is bone-conducted speech, the voice activity detector is used to stop the adaptation process. If the voice activity detector stops adapting when there is speech from a third party, the adaptation will stop unnecessarily, ultimately slowing down the adaptation.

[0194] Figure 5 The ERR (AE) and FF (AM) signal paths relative to the ambient noise are shown. ERR lags behind the FF microphone. For an ambient noise source, AE is delayed relative to AM due to the acoustic propagation delay.

[0195] Figure 6ERR(BE) and FF(BM) signal paths for bone-conducted voice sounds are shown. The ERR steers the FF microphone. If bone-conducted voice is transmitted via the ear canal EC, the direction of the voice signal is opposite to that of the ambient noise, and the FF microphone now lags behind the error microphone, resulting in a different phase response. The bone-conducted part of the voice is typically tonal, so the overall phase response to the combined signal of ambient noise and voice varies greatly according to frequency. This results in a frequency-phase difference between the two microphones across the peaks and valleys.

[0196] Figure 7 Shows the frequency-phase response of an ERR / FF transfer function with noise cancellation and the voice exhibiting peaks, based on a bone-conducted voice signal which typically contains a fundamental and harmonics. This frequency-dependent deviation in the phase difference is used to detect whether there is voice for a first operating mode.

[0197] It is noted that the airborne part of the voice signal behaves like ambient noise and does not cause a different phase response from the ambient noise, so this does not pose a problem. It is also noted that the transfer function of the bone-conducted voice propagating out of the ear varies greatly from person to person, so any metric used to detect the peaks in this response needs to simply detect "peaks" rather than a specific transfer function. Additionally, the phase response when there is no voice will vary due to leakage and ANC filter characteristics (FB and FF).

[0198] Not all voice signals show distinct harmonics, so detecting the harmonic relationship in the peaks may not yield a reliable method.

[0199] The advantage of detecting these peaks is that it only detects the bone-conducted voice of the headphone user, rather than the airborne voice path or voice from a third party. In cases where voice can interfere with adaptive noise-canceling headphones, the voice activity detector must pause the process. Detecting only the user's voice and not third-party voice signals ensures that adaptation is stopped less frequently.

[0200] Figure 8A and Figure 8B Shows an ANC performance plot, such as the feedforward target and ANC performance. In an ANC headphone FF system, when the ANC tends to be good (i.e., when the FF filter closely matches the acoustics (feedforward target)), the ANC performance can be shown as peaks and valleys. Figure 8A The graph g1 in Figure 8B shows an ANC process with worse ANC performance than the graph g2 below.

[0201] For good ANC ( Figure 8BIn the curve graph g2), the filter should match the acoustic amplitude and phase very closely. Small frequency-dependent amplitude and phase variations in the acoustic response mean that the filter intersects the acoustic response at several points, resulting in very different ANCs in adjacent frequency bands.

[0202] This means that compared to the FF signal, the error signal ERR will peak, and when the ANC approaches good performance, it will erroneously report the presence of speech. Therefore, when the solution results in suboptimal ANC, the first operating mode will erroneously detect speech and stop adaptation. Thus, when the detection parameter (in this case the ANC performance is below the threshold), the VAD switches to the second operating mode. The ANC performance is approximately the ratio of the error microphone energy to the FF microphone energy. In the case of playing music from a device, the operation is to remove the music from the error microphone signal. In cases where this removal of music is not effective enough, the ANC approximation is calculated as follows:

[0203]

[0204] where all values represent energy levels, ERR is the signal at the error microphone, FF is the signal at the FF microphone, and MUS is the sound signal or music signal.

[0205] The second operating mode only analyzes the signal at the error microphone. In this case, it monitors the error signal ERR and triggers the state of speech activity when a tone is detected. This method of detecting speech is no longer triggered only for the user's speech and will also be erroneously triggered if the ambient noise is a particular tone. This means that for high-pitched ambient noise sources, adaptation cannot exceed the ANC threshold. The ANC threshold is typically about 20 dB, although this is considered acceptable.

[0206] Figure 9 Shows the operating modes for rapid speech detection. The first two processes (referred to herein as "slow" processes) can operate in the frequency domain or be delayed by a time averaging process, and may not be able to stop adaptation quickly enough. The third process (referred to herein as the "fast" process) operates in the time domain to detect a sudden increase in the energy at the error microphone relative to the FF microphone. That is, it detects a sudden drop in the ANC performance approximation caused by speech.

[0207] The fast process is calculated as Figure 9 shown. The energy ratio between the two microphones (ERR / FF) is calculated. This ANC performance approximation energy is calculated over short and long time periods. Thus, if the ANC performance suddenly decreases, the difference between the short-term energy and the long-term energy (A) will become positive, which is usually the case when speech is present.

[0208] If the start of the sound is gradual, the slow process is considered fast enough to react appropriately.

[0209] In an adaptive ANC headset, a sudden drop in ANC performance can also be the result of a rapid change in the acoustic load around the headset, such as suddenly pushing the earbud into the ear. Before the system has time to fully adapt, the error energy will increase relative to the FF signal, which may trigger the fast process. In this case, the action can be to pause the adaptation due to the fear of the presence of speech, thus delaying the adaptation of the earbud. To correct this, the ratio (B) of the short-term energy to the long-term energy of the FF or noise signal is also monitored. If the ambient noise suddenly increases, this value is higher than 1.

[0210] Due to the air-borne speech path, this always happens when there is speech.

[0211] Therefore, applying simple logic to this arrangement can set the possible states of speech:

[0212] If A > threshold_l and B > threshold_2:

[0213] Sound = possible

[0214] This can be used as a highly aggressive VAD for adaptive ANC, because it can quickly pause the adaptation process assuming the presence of speech, and then rely on slower and more accurate metrics to re-enable the adaptation.

[0215] It is obvious to those skilled in the art that the subtraction and division stages x and y can be subtraction or division, and produce comparable functions.

[0216] Figure 10 A flowchart showing the possible operating modes of the voice activity detector is shown. The voice activity detector can operate in three or four modes:

[0217] 1. The fast process sets the possible state of speech and waits for the result from the slow process.

[0218] 2. If the ANC performance is higher than the ANC threshold, the phase difference between the two microphones is considered

[0219] 3. If the ANC performance is lower than the ANC threshold, the error microphone tone is considered.

[0220] The VAD will mainly operate in modes 1 and 2, thus providing a VAD that is sensitive to bone-conducted speech.

[0221] In the fourth mode, when music is active, VAD can enter the second mode or a combination of the first and second modes. In the case of playing music, the phase detection metric may often return false positives that are unacceptable. In this case, the logic is changed so that the tone and phase difference are monitored for voice conditions.

[0222] Both are highly likely to be triggered by voice, but much less likely to be triggered by music.

[0223] There are multiple ways to detect the peaks and tones of modes 2 and 3, and these methods have different advantages. Some examples are discussed here, but alternative peak detection and tone methods not disclosed here can be used.

[0224] In the embodiments discussed above, ANC performance parameters have been used. For example, this parameter can be defined as the ratio of ERR and FF. However, other definitions are possible such that generally detection parameters can be considered. As an example, an alternative way to monitor ANC performance (in an adaptive system) can be to look at the gradient of the adaptive parameters. When adaptation is successful, the adaptive parameters change more slowly, and thus the gradients of these parameters flatten.

Claims

1. A signal processing method for voice activity detection in an ear-worn playback device (HP), the ear-worn playback device including a speaker (SP), an error microphone (FB_MIC) that mainly senses the sound output from the speaker (SP) and also senses the surrounding ambient sound, and a feedforward microphone (FF_MIC) that mainly senses the surrounding ambient sound, the method comprises the following steps: using a voice activity detector: - recording a feedforward signal (FF) from the feedforward microphone (FF_MIC), - recording an error signal (ERR) from the error microphone (FB_MIC), - determining at least one detection parameter based on the feedforward signal (FF) and the error signal (ERR), and - monitoring the at least one detection parameter and setting a voice activity state based on the at least one detection parameter; wherein, the detection parameter: - is based on the ratio of the feedforward signal to the error signal, the method further comprises the steps of, using the voice activity detector: - determining one of the following voice activity states: false, true or likely, - the voice activity state being equal to true indicates that voice is detected, and - the voice activity state being equal to false indicates that no voice is detected, the method further comprises the steps of, using the voice activity detector: - controlling an adaptive noise cancellation controller based on the voice activity state, wherein the adaptive noise cancellation controller is coupled to the feedforward microphone and the error microphone, the adaptive noise cancellation controller performs noise cancellation processing based on the feedforward signal and / or the error signal, and by using a filter coupled to the feedforward microphone and the speaker, the filter having a filter transfer function determined by the noise cancellation processing, the method is characterized in that it further comprises the following steps: - when the detection parameter is greater than a first threshold, using the voice activity detector in a first operation mode, and when the detection parameter is less than the first threshold, using the voice activity detector in a second operation mode, - in the first operation mode, analyzing the phase difference between the feedforward signal and the error signal, and - setting the voice activity state based on the analyzed phase difference, in the second operation mode: - analyzing the pitch level of the error signal, and - setting the voice activity state based on the analyzed pitch level, the method further comprises the steps: - in a fourth operation mode, using the voice activity detector to determine whether a sound signal is active, - if the sound signal is active and if the detection parameter is less than the first threshold, using the voice activity detector to enter the second operation mode, and - if the sound signal is active and if the detection parameter exceeds the first threshold, using the voice activity detector to enter a combined first operation mode and second operation mode, the combined first operation mode and second operation mode comprising, using the voice activity detector, setting the voice activity state based on both the analyzed phase difference and pitch level.

2. An audio system for an ear-worn playback device (HP), comprising: - a speaker (SP), - an error microphone (FB_MIC) that senses sound output from the speaker and ambient sound, and - a feedforward microphone (FF_MIC) that mainly senses ambient sound, wherein the audio system includes a voice activity detector (VAD) configured to: - record a feedforward signal (FF) from the feedforward microphone (FF_MIC), - record an error signal (ERR) from the error microphone (FB_MIC), - determine at least one detection parameter based on the feedforward signal (FF) and the error signal (ERR), and - monitor the at least one detection parameter and set a voice activity state based on the at least one detection parameter, - an adaptive noise cancellation controller (ANCC) coupled to the feedforward microphone (FF_MIC) and the error microphone (FB_MIC), the adaptive noise cancellation controller (ANCC) configured to perform noise cancellation processing based on the feedforward signal (FF) and / or the error signal (ERR), - a filter (FL) coupled to the feedforward microphone (FF_MIC) and the speaker (SP), the filter having a filter transfer function (F) determined by the noise cancellation processing, wherein the at least one detection parameter: - is based on a ratio of the feedforward signal (FF) to the error signal (ERR), - is a phase difference between the error signal and the feedforward signal, or - is also based on a sound signal (MUS), and wherein: - the voice activity detector process determines one of the following voice activity states: false, true, or likely, - the voice activity state equal to true indicates that speech is detected, and - the voice activity state equal to false indicates that no speech is detected, and / or - the voice activity detector (VAD) controls the adaptive noise cancellation controller (ANCC) according to the voice activity state, and wherein, in a first operation mode, the voice activity detector (VAD): - analyzes a phase difference between the feedforward signal (FF) and the error signal (ERR), and - sets the voice activity state according to the analyzed phase difference, and / or - enters the first operation mode when the detection parameter is greater than a first threshold, and wherein, in a second operation mode, the voice activity detector (VAD): - analyzes a pitch level of the error signal (ERR), and - sets the voice activity state according to the analyzed pitch level, and / or - enters the second operation mode when the detection parameter is less than the first threshold, and wherein, in a fourth operation mode, the voice activity detector (VAD): - determines whether the sound signal (MUS) is active, - if no sound signal (MUS) is active, enters the first operation mode or the second operation mode, - If the voice signal (MUS) is active and if the detection parameter is less than the first threshold, enter the second operating mode, and - If the voice signal (MUS) is active and if the detection parameter exceeds the first threshold, enter the combined first and second operating modes.

3. The audio system according to claim 2, wherein, the noise cancellation processing includes feedforward noise cancellation processing or feedback noise cancellation processing or both feedforward noise cancellation processing and feedback noise cancellation processing.

4. The audio system according to claim 2, wherein, the control of the adaptive noise cancellation controller (ANCC) includes: - Pausing the adaptation of the noise cancellation signal when the voice activity state is set to true and / or probable, and - Continuing the adaptation of the noise cancellation signal when the voice activity state is set to false.

5. The audio system according to claim 2, wherein - Comparing the analyzed phase difference with the expected phase difference, and - Setting the voice activity state to false when the analyzed phase difference is less than the expected phase difference, otherwise setting it to true.

6. The audio system according to claim 2, wherein - Comparing the analyzed pitch level with the expected pitch level, and - Setting the voice activity state to false when the analyzed pitch level is less than the expected pitch level, otherwise setting it to true.

7. The audio system according to claim 2, wherein, in a third operating mode capable of operating independently of the first and second operating modes, the voice activity detector (VAD) is configured to: - Monitor the detection parameter during a first time period and represent it as a short-term parameter, and monitor the detection parameter during a second time period and represent it as a long-term parameter, wherein the first time period is shorter in time than the second time period, - Combine the short-term parameter and the long-term parameter to obtain a combined detection parameter, and - Set the voice activity state according to the combined detection parameter.

8. The audio system according to claim 7, wherein, in the third operating mode: - The short-term parameter and the long-term parameter are equivalent to energy levels, and - When the change in the relative energy level exceeds a second threshold, set the voice activity state to probable.

9. The audio system according to claim 2, wherein, in the combined first and second operating modes, the voice activity detector (VAD): - Analyze the pitch level of the error signal (ERR) and analyze the phase difference between the feedforward signal (FF) and the error signal (ERR), and - Set the voice activity state according to the analyzed phase difference and the analyzed pitch level, and / or in the combined first and second operating modes: - Compare the analyzed pitch level with the expected pitch level and compare the analyzed phase difference with the expected phase difference, - When the analyzed pitch level is less than the expected pitch level and the analyzed phase difference is less than the expected phase difference, set the speech activity state to false, and - When the analyzed pitch level exceeds the expected pitch level and the analyzed phase difference exceeds the expected phase difference, set the speech activity state to true.

Citation Information

Patent Citations

  • Audio classifier for half duplex communication

    US20020165718A1

  • Apparatus and method for detecting voice end point

    US20080095384A1

  • Hearing assistance system with own voice detection

    US20100260364A1

  • Noise cancellation system with gain control based on noise level

    US20100266137A1

  • Systems, methods, devices, apparatus, and computer program products for audio equalization

    US20110293103A1