Voice activity detection

CN115735362BActive Publication Date: 2026-08-21BOSE CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202180045895.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-29
Filing Date
2021-04-23
Publication Date
2026-08-21
Estimated Expiration
2041-04-23

Smart Images

  • Figure CN115735362B_ABST
    Figure CN115735362B_ABST
Patent Text Reader

Abstract

A headset that can detect voice activity of a user, the headset comprising: an internal microphone that generates an internal microphone signal; an external microphone that generates an external microphone signal, wherein the internal microphone and the external microphone are positioned such that, when the headset is worn by a user, the internal microphone is disposed closer to the user's head; and a voice activity detector that determines a sign of a phase difference between the internal microphone signal and the external microphone signal, and generates a voice activity detection signal that is representative of voice activity of a user when the sign of the phase difference indicates that the external microphone receives an audio signal after the internal microphone receives the audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Patent Application Serial No. 16 / 862,126, filed April 29, 2020, entitled “Voice Activity Detection,” the entire disclosure of which is incorporated herein by reference. Background Technology

[0003] This disclosure relates in general to speech activity detection. Various examples involve detecting a user's speech based on the phase difference between an internal and external microphone in a headset. Summary of the Invention

[0004] All examples and features mentioned below can be combined in any technically possible way.

[0005] According to one aspect, the headset includes an internal microphone that generates an internal microphone signal; an external microphone that generates an external microphone signal, wherein the internal and external microphones are positioned such that when the headset is worn by a user, the internal microphone is positioned closer to the user's head; and a voice activity detector configured to determine the sign of a phase difference between the internal and external microphone signals, and to generate a voice activity detection signal representing the user's voice activity when the sign of the phase difference indicates that the external microphone receives an audio signal after the internal microphone receives an audio signal.

[0006] In one example, the voice activity detector is further configured to convert an internal microphone signal into a frequency domain internal microphone signal including at least a first internal microphone signal phase at a first frequency, and to convert an external microphone signal into a frequency domain external microphone signal including at least a first external microphone signal phase at a first frequency, wherein the sign of the phase difference between the internal microphone signal and the external microphone is determined according to the sign of the difference between the first internal microphone signal phase and the first external microphone signal phase.

[0007] In one example, the frequency domain internal microphone signal also includes a second internal microphone signal phase at a second frequency, and the frequency domain external microphone signal also includes a second external microphone signal phase at a second frequency, wherein the sign of the phase difference between the internal microphone signal and the external microphone signal is further determined based on the sign of the difference between the second internal microphone signal phase and the second external microphone signal phase.

[0008] In one example, the sign of the phase difference is the sign of the time-domain product of the internal microphone signal and the external microphone signal.

[0009] In one example, a voice activity detection signal representing the user's voice activity is generated only if the noise present in the external microphone signal is below a threshold.

[0010] In one example, noise present in the external microphone is determined based on a measure of the similarity or linear relationship between the internal and external microphone signals.

[0011] In one example, the measure of a linear relationship is coherence.

[0012] In one example, the headset also includes an active noise canceller configured to generate a noise cancellation signal, the active noise canceller being configured to perform at least one of the following in response to generating a voice activity detection signal representing a user's voice activity: interrupting or minimizing the magnitude of the noise cancellation signal, and starting to generate a transparency signal or increasing the magnitude of the transparency signal.

[0013] In one example, the headset also includes an audio equalizer configured to receive an audio signal input and generate an audio signal output, the audio equalizer interrupting or minimizing the audio signal output by a certain amount in response to generating a voice activity detection signal representing the user's voice activity.

[0014] In one example, headphones are one of the following: headphones, earbuds, hearing aids, or mobile devices.

[0015] According to another aspect, a method for detecting a user's voice activity includes the following steps: providing a headset having an internal microphone that generates an internal microphone signal and an external microphone that generates an external microphone signal, wherein the internal and external microphones are positioned such that when the headset is worn by a user, the internal microphone is positioned closer to the user's head; determining the sign of a phase difference between the internal and external microphone signals; and generating a voice activity detection signal representing the user's voice activity when the sign of the phase difference indicates that an audio signal is received by the external microphone after the internal microphone has received an audio signal.

[0016] In one example, the method further includes the steps of: converting an internal microphone signal into a frequency domain internal microphone signal, the frequency domain internal microphone signal including at least a first internal microphone signal phase at a first frequency; and converting an external microphone signal into a frequency domain external microphone signal, the frequency domain external microphone signal including at least a first external microphone signal phase at a first frequency, wherein the sign of the phase difference between the internal microphone signal and the external microphone is determined according to the sign of the difference between the first internal microphone signal phase and the first external microphone signal phase.

[0017] In one example, the frequency domain internal microphone signal also includes a second internal microphone signal phase at a second frequency, and the frequency domain external microphone signal also includes a second external microphone signal phase at a second frequency, wherein the sign of the phase difference between the internal microphone signal and the external microphone signal is further determined based on the sign of the difference between the second internal microphone signal phase and the second external microphone signal phase.

[0018] In one example, the sign of the phase difference is the sign of the time-domain product of the internal microphone signal and the external microphone signal.

[0019] In one example, a voice activity detection signal representing the user's voice activity is generated only if the noise present in the external microphone signal is below a threshold.

[0020] In one example, noise present in the external microphone is determined based on a measure of the similarity or linear relationship between the internal and external microphone signals.

[0021] In one example, the measure of a linear relationship is coherence.

[0022] In one example, the method further includes the steps of performing at least one of the following in response to generating a voice activity detection signal representing the user's voice activity: interrupting active noise cancellation or minimizing the magnitude of active noise cancellation, and starting to generate a transparent signal or increasing the magnitude of the transparent signal.

[0023] In one example, the method further includes the step of interrupting or minimizing the generation of an audio signal in response to generating a voice activity detection signal representing the user's voice activity.

[0024] In one example, the internal and external microphones are set on one of the following: headphones, earbuds, hearing aids, or mobile devices.

[0025] Details of one or more specific embodiments are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the specification, drawings, and claims. Attached Figure Description

[0026] In the accompanying drawings, similar reference numerals generally refer to the same parts in all different views. Furthermore, the drawings are not necessarily drawn to scale, and the focus is usually on illustrating the principles governing the various aspects.

[0027] Figure 1 A perspective view is depicted based on an example of a headset with voice activity detection using an internal and external microphone.

[0028] Figure 2A perspective view is depicted based on an example of a headset with voice activity detection using an internal and external microphone.

[0029] Figure 3 A block diagram of a speech activity detector based on an example is depicted.

[0030] Figure 4 A graph depicting the phase difference across frequencies between the internal and external microphones is presented.

[0031] Figure 5 A block diagram of a speech activity detector and an active noise canceller based on an example is depicted.

[0032] Figure 6 A block diagram of a speech activity detector and an audio equalizer based on an example is depicted.

[0033] Figure 7A A flowchart illustrating voice activity detection using an internal and external microphone, based on an example, is provided.

[0034] Figure 7B A flowchart illustrating voice activity detection using an internal and external microphone, based on an example, is provided.

[0035] Figure 7C A flowchart illustrating voice activity detection using an internal and external microphone, based on an example, is provided.

[0036] Figure 7D A flowchart illustrating voice activity detection using an internal and external microphone, based on an example, is provided. Detailed Implementation

[0037] Generally, it is undesirable to generate active noise cancellation signals that eliminate ambient noise (rather than, for example, the user's own voice), or to generate audio output in headphones worn by a user who is speaking or otherwise engaging in a conversation. Therefore, it is desirable to detect the user's voice and, upon detection, interrupt any audio output from the headphones that would disrupt or interfere with the user's conversation. The various examples disclosed herein describe detecting a user's voice activity by comparing the phases of two microphones positioned on the headphones.

[0038] exist Figure 1 and Figure 2 Exemplary headsets 100 and 200 with voice activity detection are shown. First turn to Figure 1 The headphones 100 have a connection to the left earpiece 104. L and right earpiece 104 RA pair of over-ear headphones with headband 102. Left earpiece 104. L Including internal microphone 106 L and external microphone 108 L The left earpiece also includes a transducer 110 for converting noise-cancelled signals or any other input audio signals. L (That is, the speaker). Similarly, the right earpiece 104 R Including internal microphone 106 R 108 external microphones R and transducer 110 R The 200 headphones are a pair of in-ear headphones, including the left earpiece 204. L and right earpiece 204 R Extending from it is a bushing 202. Similar to headphones 100, the earpiece 204... L- and 204 R Each includes an internal microphone 106 L 106 R 108 external microphones L 108 R and transducer 110 L 110 R .

[0039] In most examples, the internal microphone 106 is located on the inner surface of the headset, such as within the ear cup of the headset (e.g., as shown in the image). Figure 1 (as shown) or positioned inside the user's ear (e.g., as shown) Figure 2 As shown), the external microphone 108 is located on the outer surface of the headset, such as on the outside of the earpiece (e.g., as shown). Figure 1 and Figure 2 (As shown). However, it is only necessary to position the internal microphone 106 closer to the user's head than at least one corresponding external microphone 108, such that the user's voice signal (e.g., converted by bone, tissue, air, or other medium) reaches the internal microphone 106 before it reaches the corresponding external microphone 108.

[0040] Although a single internal microphone 106 and an external microphone 108 are shown on each earpiece 104, 204, any number of internal microphones 106 and external microphones 108 can be used. Furthermore, the number of internal microphones 106 and external microphones 108 need not be the same. For example, in some examples, each earpiece 104, 204 may include two internal microphones 106 and three external microphones 108.

[0041] For the purposes of this disclosure, a headset is any device that is worn by a user or otherwise held against a user's head and includes a transducer for playing audio signals such as noise cancellation signals or audio signals. In various examples, a headset may include headphones, earbuds, hearing aids, or mobile devices.

[0042] Each headset 100, 200 includes a voice activity detector 300, which in... Figure 3 The block diagram illustrates this. The voice activity detector 300 determines when a user wearing or otherwise using the headset is speaking based on the sign of the phase difference between the signals output from the internal microphone 106 and the external microphone 108. In various examples, the voice activity detector 300 may be implemented in a controller such as a microcontroller, which includes a processor and a non-transitory storage medium storing program code that performs the various functions of the voice activity detector 300 described herein when executed by the processor. Alternatively, the voice activity detector 300 may be implemented in hardware, such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA). In yet another example, the voice activity detector may be implemented as a combination of hardware, firmware, and / or software.

[0043] like Figure 3 As shown, the voice activity detector 300 receives an internal microphone signal u from the internal microphone 106. 内部 and the external microphone signal u from external microphone 108 外部 .Although Figure 3 Only one internal microphone signal u received from a single internal microphone 106 is shown. 内部 and an external microphone signal u received from a single external microphone 108 外部 However, it should be understood that in other examples, the voice activity detector 300 can receive and use any number of internal microphone signals. 内部 and external microphone signal u 外部 .

[0044] As described above, the voice activity detector 300 determines the internal microphone signal u 内部 and external microphone signal u 外部 The sign of the phase difference between the internal and external microphone signals is used to detect the user's speech activity. The phase difference between the internal and external microphone signals indicates the directionality of the input audio signal. This is because the audio signal will be delayed as it travels from the audio source to one microphone and then to another. For example, if the audio signal originates from a point A closer to the internal microphone 106 (e.g., from user speech activity arising from tissue and bone transitions in the user's head), the audio signal will travel a distance d. A1To reach the internal microphone 106, but the travel distance d A2 (This distance is longer than distance d) A1 The audio signal originating from point A will first reach the internal microphone 106, and then the external microphone 108. Conversely, if the audio signal originates from point B, which is closer to the external microphone 108 (e.g., from an audio source farther from the user), the audio signal will travel a distance d. B1 To reach the external microphone 108, but the travel distance d B2 (This distance is longer than distance d) B1 The audio signal originating from point B will first reach the external microphone 108, and then the internal microphone 106. The delay between the audio signals arriving at the internal microphone 106 and the external microphone 108 will be determined by the distance between the two microphones. From the signal's perspective, this delay will be represented by the internal microphone signal u. 内部 and external microphone signal u 外部 The phase difference between them.

[0045] The relative delay determines the sign of the phase difference between the internal and external microphone signals. Therefore, when the audio signal originates from outside the headphones, the phase difference will have one sign (e.g., positive); however, when the audio signal originates from inside the headphones, the phase difference will have the opposite sign (e.g., negative). Thus, the internal microphone signal u 内部 and external microphone signal u 外部 The phase difference between them indicates the user's voice activity.

[0046] For an audio signal originating from a given point (the user's voice activity or an external source), whether the phase difference is positive or negative depends on whether the phase difference originates from the internal microphone signal u. 内部 Or from the external microphone signal u 外部 Measured. For example, from the internal microphone signal u. 内部 external microphone signal u 外部 The measured 90° phase difference will be from the external microphone signal u 外部 to internal microphone signal u 内部 The measured -90° phase difference. Therefore, for the purposes of this disclosure, the internal microphone signal u can be obtained. 内部 external microphone signal u 外部 Or from an external microphone signal u 外部 to internal microphone signal u 内部Measure the phase difference. (Only a 90° phase difference is provided as an example. It should be understood that the magnitude of the phase difference will depend on the distance between the internal microphone 106 and the external microphone 108, as well as the frequency at which the phase difference is measured.)

[0047] Phase difference can be measured in any suitable manner. In the first example, the phase difference can be measured by transforming the internal and external microphone signals to the frequency domain and comparing the phases of the microphone signals at at least one representative frequency. For example, the internal and external microphone signals can be processed using a Discrete Fourier Transform (DFT) to generate multiple frequency windows, each including phase information of the associated microphone signal at the corresponding frequency. Then, a microphone signal derived from the DFT at at least one representative frequency (e.g., the internal microphone signal u) is... 内部 The phase information of the microphone signal is compared with that of another microphone signal at the same or different representative frequencies (e.g., external microphone signal u). 外部 The phase information of ) is compared. In Figure 4 The figure illustrates an example of the results of such a transformation, showing the signals from 12 internal microphones spanning a frequency band from 100Hz to 1000Hz when the user is speaking (labeled as speech) and when the user is not speaking (labeled as external noise). 内部 and external microphone signal u 外部 The graph shows the phase difference between the two frequencies. From approximately 250Hz to 600Hz, the phase difference varies between approximately 180° and 0°; however, when the user is not speaking, the phase difference within the same frequency band ranges from approximately -20° to -90°. In this example, the internal microphone signal u is at any frequency within the 250Hz to 600Hz range. 内部 and external microphone signal u 外部 The positive phase difference will accurately match the user's voice activity.

[0048] While DFT typically generates phase information across multiple frequency windows, in one example, the phase at only a single representative frequency can be determined and used to determine the phase difference. This single representative frequency could be, for example, the average center frequency of the human voice conducted through bone / tissue. For instance, a typical female voice generates an acoustic excitation from 200 Hz to 1000 Hz at an internal microphone, so a phase difference at a center frequency of 600 Hz could be used. Alternatively, representative frequencies that typically represent the phase difference symbols corresponding to the user's speech can be determined empirically.

[0049] However, a phase difference at a single frequency is not necessarily suitable for determining the phase difference that will reliably match the user's speech, because the speech quality and frequency range of the user's speech will vary from user to user. Figure 3 As shown, the sign of the phase difference will vary across frequencies, therefore the sign of the phase difference used for speech activity detection can be determined from multiple different phase differences acquired at various different frequencies. Thus, in an alternative example, the phase at multiple frequency windows can be used to determine the internal microphone signal u. 内部 and external microphone signal u 输出 The phase difference. Any number of methods can be used to determine the phase difference from phases at multiple frequencies. For example, the phase difference can be determined based on the signs of most of the phase differences at multiple frequencies. Thus, for five phase differences p1 to p5, each acquired at corresponding representative frequencies f1 to f5, the phase difference used to determine whether a user is speaking can be determined to be positive if three or more of these five phase differences are positive. However, if three or more of the five phase differences are negative, the phase difference can be determined to be negative. Alternatively, for a phase difference to be determined to be positive, some threshold number of phase differences must be positive. For example, if two of the five phase differences are positive, or if one of the five phase differences is positive, the phase difference can be determined to be positive. In yet another example, the sign of the median phase difference among multiple phase differences can be used as the phase difference sign to determine whether a user is speaking. In the case of using phase differences of multiple frequency values ​​to determine whether a user is speaking, the frequency window used can be continuous, or alternatively, the frequency window used can be divided into one or more frequency windows.

[0050] Although the DFT has been discussed in this paper, any method used to determine the phase of a signal at at least one representative frequency can be used. In alternative examples, the Fast Fourier Transform (FFT) or the Discrete Cosine Transform (DCT) can be used.

[0051] In another example, the internal microphone signal u can be determined in the time domain. 内部 and external microphone signal u 外部 The phase difference between them, rather than the internal microphone signal u 内部 and external microphone signal u 外部 Transform to the frequency domain. For example, the internal microphone signal u. 内部 and external microphone signal u 外部 The sign of the phase difference between them can be determined by the internal microphone signal u. 内部 and external microphone signal u 外部 The time-domain product (e.g., internal microphone signal u) 目标 and external microphone signal u 外部 The product of one or more samples is used to determine the internal microphone signal u. If the product is positive, the internal microphone signal u can be determined. 内部 and external microphone signal u 外部The phase difference between them is positive. However, if the product is negative, the internal microphone signal u can be determined. 内部 and external microphone signal u 外部 The phase difference between them is negative. One or both of these time-domain signals can be filtered (e.g., bandpass filtered) to improve the phase estimation within a certain frequency range of interest.

[0052] In the presence of multiple internal microphones 106 and / or multiple external microphones 108, phase differences between any number and combination of internal microphones 106 and external microphones 108 can be found. For example, if a headset includes three internal microphones 106 and three external microphones 108, a phase difference between each of the three internal microphones can be found for each of the three external microphones, resulting in nine separate phase differences. Thus, the number of internal microphones 106 and external microphones 108 does not need to be symmetrical. In fact, a phase difference can be found between one internal microphone and three external microphones, resulting in three phase differences. Alternatively, the phase difference for each internal microphone can be found for only one external microphone. The only limitation is that the internal microphones 106 are positioned relative to the external microphones 108 to receive the user's voice in front of the external microphones 108.

[0053] When voice activity is detected, the voice activity detector 300 generates a voice activity detection signal. The voice activity detection signal can be a binary signal having a first value (e.g., 1) when voice activity is detected and a second value (e.g., 0) when no voice activity is detected. In an alternative example, these values ​​can be reversed (e.g., 1 when voice activity is detected and 0 when no voice activity is detected). Furthermore, the voice activity detection signal can be a signal internal to the controller and can be stored and referenced by other subsystems or modules within the headset for indicating other functions. For example, the headset's active noise cancellation system can be turned on / off based on the value of the voice activity detection signal.

[0054] The reliability of the phase difference between the internal and external microphones will be compromised in the presence of diffuse noise. For example, in a noisy environment, the internal microphone signal u 内部 The content can be combined with external microphone signals. 外部 The content is irrelevant, and therefore any measured phase difference does not indicate audio signal delay. Therefore, the voice activity detector 300 can be configured to output only a voice activity detection signal indicating the user's voice activity when the noise is below a threshold. This can be achieved by measuring the internal microphone signal u. 内部 and external microphone signal u 外部Noise can be detected by relating or similaring relationships between data points. For example, the voice activity detector 300 can measure the internal microphone signal u. 内部 and external microphone signal u 外部 The coherence between the phase difference and the phase difference (coherence is a measure of linearity). If the coherence exceeds a threshold (e.g., 0.5), it can be determined that the measured phase difference will detect the internal microphone signal u. 内部 and external microphone signal u 外部 The delay between them. Alternatively, any measure of relation or similarity can be used. For example, correlation can be used instead of coherence to determine the internal microphone signal u. 内部 and external microphone signal u 外部 The similarity.

[0055] While the internal microphone 106 and external microphone 108 can be dedicated voice activity detection microphones, in alternative examples, the internal and external microphones can be used for dual purposes, such as as input to the active noise canceller 500. Figure 5 As shown in the diagram. In operation, the active noise canceller 500 generates a noise cancellation signal c from the transducer 110. 输出 This noise cancellation signal is out of phase with and disruptively interferes with ambient noise, thereby eliminating or reducing the noise perceived by the user. Such active noise cancellers are generally known, and any suitable active noise canceller can be used in headphones. Internal microphone signal u 内部 and external microphone signal u 外部 These can be used as feedback and feedforward signals, respectively. Alternatively, a separate microphone signal can be used for noise cancellation purposes.

[0056] Similarly, the active noise canceller 500 can provide a transparent signal h 输出 For the purposes of this disclosure, active hear-through alters the active noise cancellation parameters of the headphones, allowing the user to hear some or all of the ambient sounds in the environment. The purpose of active hear-through is to allow the user to hear ambient sounds as if they were not wearing the headphones at all, and to further control their volume level. In one example, the hear-through signal h is provided in the following manner. 输出 The method employs one or more feedforward microphones (e.g., external microphone 108) to detect ambient sound, and adjusts the ANR filter for at least the feedforward noise cancellation loop to allow a controlled amount of ambient sound to pass through the earpiece with a cancellation different from what would otherwise be applied (i.e., in normal noise cancellation operation). Such an active listening method is described in US 9,949,017, entitled "Controlling ambient sound volume," but any suitable listening method may be used, the entire contents of which are incorporated herein by reference.

[0057] Noise cancellation signals can be generated in a way that does not interfere with users participating in the session. 输出 Generally, users do not want noise cancellation that reduces ambient noise while speaking or otherwise engaging in conversation. Therefore, the active noise canceller 500 can receive a voice activity detection signal v 输出 And determine whether a noise cancellation signal c is generated. 生成 As a result, for example, once the active noise canceller 500 receives a voice activity detection signal v indicating that the user is speaking. 输出 (for example, v) 输出 If the value is 1, the generation of the noise cancellation signal c can be interrupted while the user is speaking or for a period of time after the user has finished speaking. 输出 Or its magnitude can be reduced. (Generally speaking, a user who is speaking is participating in the conversation, therefore listening to the response, and may speak again soon.) Similarly, in another example, or in the same example, the listening signal h can be started while the user is speaking or some time after the user has finished speaking. 输出 Or increase its value. One or both of the following measures can be taken to allow users to participate in the conversation more naturally without interfering with active noise cancellation: reduce the noise cancellation signal c 输出 The magnitude of the noise cancellation signal or interruption, or the start of listening to the signal h. 输出 Or increase the value of the transparent signal.

[0058] Similarly, such as Figure 6 As shown, the input audio signal a can be paused. 输入 For example, music playback. Similar to noise cancellation signals, it's not necessarily desirable to play music while the user is speaking or having a conversation. The audio equalizer 600 receives input audio signals from external sources such as mobile devices or computers, or from local storage devices. 输入 And generate output a to transducer 110. 输出 Generally speaking, an audio equalizer includes one or more filters used to adjust the... 输入 And produce a 输出 The latter is converted into an audio signal by transducer 110. Audio equalizer 600 can be further configured to route the signal to multiple transducers 110. In one example, audio equalizer 600 receives v from voice activity detector 300. 输出 In response, the output audio signal a is paused. 输出 Or minimize its magnitude. For example, once the speech activity detection signal v 输出 When a user's voice activity is detected, the audio equalizer can adjust the output audio signal a. 输出Fade out until the user finishes speaking. Additionally, the audio equalizer can fade out the audio signal after the user finishes speaking. 输出 There is a delay before the fade-in occurs.

[0059] Figure 5 and Figure 6 The active noise canceller 500 and audio equalizer 600 may each be implemented in a controller, such as a microcontroller, which includes a processor and a non-transitory storage medium storing program code that, when executed by the processor, performs the various functions of the active noise canceller 500 and audio equalizer 600 described herein. The active noise canceller 500 and audio equalizer 600 may be implemented on the same controller or on separate controllers. Similarly, one or both of the active noise canceller 500 and audio equalizer 600 may be implemented on the same controller as the voice activity detector 300. Alternatively, the active noise canceller 500 and audio equalizer 600 may be implemented in hardware, such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA). In yet another example, the active noise canceller 500 and audio equalizer 600 may each be implemented as a combination of hardware, firmware, and / or software.

[0060] Figure 7A A flowchart is shown of a method 700 for detecting a user's voice activity, performed by a headset such as headset 100 or headset 200. The headset of method 700 includes at least one internal microphone and at least one external microphone, which are positioned such that when the headset is worn by a user, the internal microphone is positioned closer to the user's head than the external microphone, such that the internal microphone receives the user's voice signal before the external microphone. For example, the steps of method 700 can be implemented as steps defined in program code stored on a non-transitory storage medium and executed by a processor of a controller disposed within the headset. Alternatively, the method steps can be executed by the headset using a combination of hardware, firmware, and / or software.

[0061] At step 702, internal microphone signals and external microphone signals are received. Although only two microphone signals are described here, any number of internal and external microphone signals can be received. In fact, it should be understood that the steps of method 700 can be repeated for any combination of multiple internal and external microphone signals.

[0062] At step 704, the sign of the phase difference between the internal and external microphones is determined. This step may require first transforming the internal and external microphone signals to the frequency domain, such as using a DFT, and finding the phase difference between the phases of the internal and external microphone signals at at least one representative frequency. Alternatively, the phase difference can be determined from multiple phase differences calculated at multiple frequencies. In yet another example, the phase difference can be found in the time domain. For example, the sign of the phase difference can be determined by finding the sign of the product of one or more samples of the internal and external microphone signals. One or both of these signals can be filtered (e.g., bandpass filtered) to improve the phase estimation within a certain frequency range of interest.

[0063] At step 706, the sign of the phase difference determined in step 704 is used to detect the user's voice activity. Step 706 is therefore represented as a decision box that queries whether the sign of the phase difference between the internal and external microphones indicates that the internal microphone received the audio signal first (the sign can be positive or negative, depending on how the phase difference is calculated). If the sign indicates that the internal microphone received the audio signal before the external microphone, a voice activity detection signal indicating the user's voice activity is generated (at step 708); if the sign indicates that the external microphone received the audio signal before the internal microphone, a voice activity signal not indicating the user's voice activity is generated (step 710). Because this is a binary determination, if the sign of the phase difference does not indicate that the internal microphone received the audio signal first, it indicates that the external microphone received the audio signal first. Therefore, the decision block can be restated to query whether the phase difference indicates that the external microphone received the audio signal first, in which case the "yes" and "no" branches are reversed.

[0064] As described above, at step 708, a voice activity detection signal indicating user voice activity is generated. Conversely, at step 710, a voice activity detection signal indicating no user voice activity is generated. The voice activity detection signal can therefore be a binary signal having a value for voice detection (e.g., 1) and a value for no voice detection (e.g., 0). Because a signal with a value of 0 is generally a signal with a value of 0V, it should be understood that, for the purposes of this disclosure, the absence of a signal can be considered as the generated signal if its absence is interpreted by another system or subsystem as indicating voice detection or no voice detection.

[0065] Figure 7BAn alternative example of method 700 is depicted, wherein step 712 occurs between steps 702 and 704. Step 712 is represented as a decision box that queries whether a measure of the linear relationship or similarity between the internal and external microphone signals exceeds a threshold. Such a measure of the linear relationship could be, for example, coherence, while a measure of the similarity could be, for example, correlation. The purpose of this step is to determine whether diffuse noise dominates the internal and external microphone signals, lacking directionality sufficient to find a meaningful phase difference between the internal and external microphone signals. In the alternative example, any method for detecting ambient noise can be used. If the measure of the linear relationship or similarity exceeds the threshold, the method proceeds to step 704, where the phase difference is found as described above. Alternatively, if the measure of the linear relationship does not exceed the threshold, the step proceeds to step 710, where a voice activity detection signal indicating the absence of user voice activity is generated. In the alternative example, this step can be performed in other parts of method 700, such as after the phase difference has been found.

[0066] Figure 7C and Figure 7D It describes some optional actions that occur after a user's voice activity is detected. Figure 7C In step 712, the noise cancellation signal output from the headset transducer is interrupted to eliminate or otherwise minimize noise perceived by the user, or to reduce its magnitude. The noise cancellation signal may be interrupted or reduced until the user's voice is no longer detected, or for a predetermined period of time after the user's voice is no longer detected. Alternatively, or in a supplement to step 712, at step 714, a transparent signal output from the headset transducer to allow the user to hear some ambient noise may be generated, or the magnitude of this signal may be increased. Therefore, after the user's voice is detected, a transparent signal may be generated, or its magnitude increased, until the user's voice is no longer detected, or for a predetermined period of time after the user's voice is no longer detected. Similarly, Figure 7D The description describes interrupting the audio signal output from the headset transducer, such as music received from a mobile device or computer, at step 716. For example, the audio output signal can be faded out after a user's voice is detected. The audio output signal can be interrupted until the user's voice is no longer detected, or for a predetermined period of time after the user's voice is no longer detected. Although Figure 7C and Figure 7D This is presented as an alternative, but in other examples, any combination of steps 712, 714, and 716 can be implemented.

[0067] The functions or portions thereof described herein, and their various modifications (hereinafter referred to as "functions") may be implemented at least in part by computer program products, such as computer programs tangibly implemented in an information carrier, such as one or more non-transitory machine-readable media or storage devices, for performing or controlling the operation of one or more data processing devices, such as programmable processors, computers, multiple computers and / or programmable logic components.

[0068] Computer programs can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. Computer programs can be deployed on a single computer, distributed across one or more sites, or executed on multiple computers interconnected via a network.

[0069] The actions associated with implementing all or part of the functionality can be performed by one or more programmable processors executing one or more computer programs to perform the functions of the calibration process. All or part of the functionality can be implemented as special-purpose logic circuitry, such as FPGAs and / or ASICs (Application-Specific Integrated Circuits).

[0070] Processors suitable for executing computer programs include, for example, both general-purpose microprocessors and special-purpose microprocessors, as well as any one or more processors in any type of digital computer. Generally, a processor receives instructions and data from read-only memory or random access memory, or both. The components of a computer include a processor for executing instructions and one or more memory devices for storing instructions and data.

[0071] While several embodiments of the invention have been described and illustrated herein, those skilled in the art will readily conceive of a variety of other means and / or structures for performing the functions described herein and / or obtaining one or more of the results and / or advantages described herein, and each of such variations and / or modifications is considered to be within the scope of the embodiments of the invention described herein. More generally, those skilled in the art will readily understand that all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications using the teachings of this invention. Those skilled in the art will recognize, or can determine, many equivalents of the specific embodiments of the invention described herein using only conventional experimentation. Therefore, it should be understood that the above embodiments are presented by way of example only, and that the embodiments of the invention may be practiced in ways other than those specifically described and claimed within the scope of the appended claims and their equivalents. The embodiments of the invention disclosed herein relate to each individual feature, system, article of manufacture, material, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles of manufacture, materials, and / or methods is included within the scope of the invention disclosed herein, provided that such features, systems, articles of manufacture, materials, and / or methods do not contradict each other.

Claims

1. A type of over-ear headphone, the over-ear headphone comprising: An internal microphone that generates an internal microphone signal; An external microphone that generates an external microphone signal, wherein the internal microphone and the external microphone are positioned such that when the headset is worn by a user, the internal microphone is positioned closer to the user's head; and A voice activity detector is configured to determine the sign of the phase difference between the internal microphone signal and the external microphone signal, and to generate a voice activity detection signal indicating the user's voice activity when the sign of the phase difference indicates that the external microphone receives the audio signal after the internal microphone receives the audio signal.

2. The headset of claim 1, wherein the voice activity detector is further configured to convert the internal microphone signal into a frequency domain internal microphone signal comprising at least a first internal microphone signal phase at a first frequency, and to convert the external microphone signal into a frequency domain external microphone signal comprising at least a first external microphone signal phase at the first frequency, wherein the sign of the phase difference between the internal microphone signal and the external microphone is determined according to the sign of the difference between the first internal microphone signal phase and the first external microphone signal phase.

3. The headset of claim 2, wherein the frequency domain internal microphone signal further includes a second internal microphone signal phase at a second frequency, and the frequency domain external microphone signal further includes a second external microphone signal phase at the second frequency, wherein the sign of the phase difference between the internal microphone signal and the external microphone is further determined according to the sign of the difference between the second internal microphone signal phase and the second external microphone signal phase.

4. The headset according to claim 1, wherein the sign of the phase difference is the sign of the time-domain product of the internal microphone signal and the external microphone signal.

5. The headset of claim 1, wherein the voice activity detection signal representing the user's voice activity is generated only when the noise present in the external microphone signal is below a threshold.

6. The headset of claim 5, wherein the noise present in the external microphone is determined based on a measure of the similarity or linear relationship between the internal microphone signal and the external microphone signal.

7. The headphones of claim 6, wherein the measure of linearity is coherence.

8. The headset of claim 1 further includes an active noise canceller configured to generate a noise cancellation signal, the active noise canceller being configured to perform at least one of the following in response to the generation of a voice activity detection signal representing the user's voice activity: interrupting the noise cancellation signal or minimizing the magnitude of the noise cancellation signal, and starting to generate a transparency signal or increasing the magnitude of the transparency signal.

9. The headset of claim 1, further comprising an audio equalizer configured to receive an audio signal input and generate an audio signal output, the audio equalizer interrupting the audio signal output or minimizing the magnitude of the audio signal output in response to the generation of a voice activity detection signal representing the user's voice activity.

10. The headphones of claim 1, wherein the headphones are one of: headphones, earbuds, hearing aids, or mobile devices.

11. A method for detecting a user's voice activity, the method comprising the following steps: A headset is provided having an internal microphone that generates an internal microphone signal and an external microphone that generates an external microphone signal, wherein the internal microphone and the external microphone are positioned such that when the headset is worn by a user, the internal microphone is positioned closer to the user's head; Determine the sign of the phase difference between the internal microphone signal and the external microphone signal; and When the sign of the phase difference indicates that the external microphone receives the audio signal after the internal microphone receives the audio signal, a voice activity detection signal indicative of the user's voice activity is generated.

12. The method of claim 11, further comprising the step of: The internal microphone signal is converted into a frequency domain internal microphone signal, the frequency domain internal microphone signal including at least a first internal microphone signal phase at a first frequency; and The external microphone signal is converted into a frequency domain external microphone signal, the frequency domain external microphone signal including at least a first external microphone signal phase at the first frequency, wherein the sign of the phase difference between the internal microphone signal and the external microphone is determined according to the sign of the difference between the first internal microphone signal phase and the first external microphone signal phase.

13. The method of claim 12, wherein the frequency domain internal microphone signal further includes a second internal microphone signal phase at a second frequency, and the frequency domain external microphone signal further includes a second external microphone signal phase at the second frequency, wherein the sign of the phase difference between the internal microphone signal and the external microphone is further determined according to the sign of the difference between the second internal microphone signal phase and the second external microphone signal phase.

14. The method of claim 11, wherein the sign of the phase difference is the sign of the time-domain product of the internal microphone signal and the external microphone signal.

15. The method of claim 11, wherein the voice activity detection signal representing the user's voice activity is generated only when the noise present in the external microphone signal is below a threshold.

16. The method of claim 15, wherein the noise present in the external microphone is determined based on a measure of the similarity or linear relationship between the internal microphone signal and the external microphone signal.

17. The method of claim 16, wherein the measure of the linear relationship is coherence.

18. The method of claim 11, further comprising the step of: In response to the generation of a voice activity detection signal representing the user's voice activity, at least one of the following is performed: interrupting active noise cancellation or minimizing the magnitude of the active noise cancellation, and starting to generate a transparent signal or increasing the magnitude of the transparent signal.

19. The method of claim 11, further comprising the step of: In response to the generation of a voice activity detection signal representing the user's voice activity, the generation of the audio signal is interrupted or minimized.

20. The method of claim 11, wherein the internal microphone and the external microphone are mounted on one of: an earphone, an earplug, a hearing aid, or a mobile device.

Citation Information

Patent Citations

  • Controlling ambient sound volume

    US9949017B2

  • Systems, methods, apparatus, and computer-readable media for multi-microphone location-selective processing

    US20120020485A1

  • Headset with hear-through mode

    US20170193978A1

  • User Voice Activity Detection Methods, Devices, Assemblies, and Components

    US20180225082A1