Voice activity detection for microphone muting
The integration of VAD and noise cancellation algorithms in voice communication systems addresses unintended mute notifications by intelligently determining when to provide mute status alerts, improving user experience by reducing distractions.
Patent Information
- Application Number
- JP2023562986
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-15
- Filing Date
- 2022-03-24
- Publication Date
- 2025-10-23
- Estimated Expiration
- 2042-03-24
AI Technical Summary
Existing voice communication systems, such as headsets, often provide unintended and disruptive mute status notifications when a user's microphone is muted, particularly due to ambient noise or speaking to nearby individuals, leading to user distraction.
Implementing a voice activity detection (VAD) algorithm with additional microphones and noise cancellation algorithms to intelligently determine when a mute status notification is necessary, considering factors like ambient noise, user conversation, and call activity, ensuring notifications are only triggered when the user intends to speak during a call.
Significantly reduces unintended mute status notifications by accurately distinguishing between user speech and ambient noise, thereby enhancing the user experience during voice communications.
Smart Images

Figure 0007759401000001 
Figure 0007759401000002 
Figure 0007759401000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of voice communications, such as two-way voice communications over a communications link, e.g., online two-way communications. Specifically, the present invention proposes a method for informing a user to mute a microphone, e.g., using one or more microphone inputs, to eliminate or reduce intrusive notifications to the user based on a voice activity detection algorithm. [Background technology]
[0002]
[0003] Headsets have many advantages for participating in online calls and conferences, but some advantages also have disadvantages. For example, it is desirable to be able to mute the headset microphone when a user has nothing to add to a call for a while. One disadvantage of this feature is that a user may forget that the microphone is muted when they want to talk on the call.
[0003] This problem is solved in some cases by detecting whether a user is speaking into a muted microphone and providing a visual or audio notification to the user. Audio notification can be advantageous because a user who is on a call will always hear the notification. However, one drawback of this feature is that a user may sometimes intentionally speak into a muted microphone.
[0004] For example, if a user is on a call and at some point wants to speak to a colleague who is physically present near the user, the user's microphone is muted and an audio notification is played when the user starts talking to the colleague, resulting in the audio notification disrupting the user rather than providing assistance.
[0005] Furthermore, another drawback is that headset microphones also pick up ambient sounds, which may include colleagues talking. This can lead to situations where a user is on a call with their microphone muted, and when their colleague is speaking, the headset microphone picks up their speech and alerts the user that they are speaking into the muted microphone, when in fact they are not.
[0006] EP2881946A1 describes a microphone mute / unmute system that detects quiet events and voice activity from audio signals at the far end and near end to determine whether an audio event is disruptive or a talker. The described system may also detect faces and movements from camera images to determine a mute or unmute indication.
[0007] US2015 / 0195411A1 describes a system for providing intelligent, automatic mute notifications that utilizes a combination of recorded characteristics and triggered timers to provide a mechanism for controlling false positives of speech while muted. Summary of the Invention [Problem to be solved by the invention]
[0008] Therefore, in accordance with the above, it is an object of the present invention to provide a method and apparatus that eliminates or reduces the problem of unintended notifications and alerts when a user is participating in a call or conversation with their microphone muted. [Means for solving the problem]
[0009] In a first aspect, the present invention provides a method of notifying a user of a muted state of a microphone system when the user speaks during a call with one or more other participants while a primary microphone of the microphone system is muted, the method comprising: - processing an output signal from the microphone system, at least the primary microphone, based on a voice activity detection (VAD) algorithm using a processor system while the primary microphone of the microphone system is muted; - determining whether speech is present based on the output of a voice activity detection algorithm; - determining whether additional conditions are met; and - providing a mute status notification to the user only when speech is present and additional conditions are determined to be met; Includes:
[0010] Preferably, the microphone system comprises a primary microphone configured to capture the user's speech and an additional microphone positioned to capture sounds around the user, and the method includes processing output signals from the primary microphone and the additional microphone to perform a noise cancellation algorithm to suppress ambient noise. The additional microphone is preferably not connected to transmit audio to the call at any time, as its role is to provide information to determine whether the primary microphone should be muted or whether a mute notification should be sent.
[0011] Such a method is beneficial because it can eliminate or at least reduce many of the problems associated with simple notification of a muted microphone, thus avoiding significant disruption in various situations during a call. This method is particularly suited to headsets configured to connect to computers, tablets, smartphones, or the like. Generally, this method can be used with various devices that have microphones intended for use in calls and that have microphone mute capabilities. Portions of the method can be advantageously implemented in a processor of a headset or other device that includes a microphone and speaker, while other portions of the method can be implemented in components involved in facilitating the call, such as a computer or a server providing the online call. In particular, a headset or other device can simply mute its primary microphone to remove notifications of unintended muting from an online call.
[0012] The present invention is based on the finding that a significant number of unintentional microphone mute notifications can be intelligently eliminated or reduced by fairly simple processing that can be implemented, for example, in a headset. Using a voice activity detection (VAD) algorithm, it is possible to ensure that only speech at a muted microphone triggers a mute status notification. By further introducing additional conditions that must be met before a mute status notification is provided to the user, user activity can be intelligently detected to significantly reduce intrusive mute status notifications, for example, when the user is speaking to a colleague physically nearby. If the VAD algorithm notices that the user is speaking into the microphone, it sends an interrupt, causing the headset to play a voice prompt, for example, "headset is muted." The mute status notification can be audible, as this form of notification is most likely to be noticed by the headset user. Preferably, the frequency of the mute status notification should be configurable, for example, by the user, to prevent the notification from being played too frequently when the headset user is physically speaking.
[0013] In particular, the use of more microphones apart from the user's primary microphone can be used to detect other people speaking in the environment. Further, other VAD algorithms can be used to detect speech captured by such additional microphones. Further, other VAD algorithms can be applied to the incoming audio signal in a call so that other participants can be detected speaking, indicating that the user may intend to speak, for example by allowing a mute status notification to the user only when no other participants are speaking in the call.
[0014] Noise cancellation (ENC or ANC) using one or more additional microphones assists the VAD algorithm and improves its ability to determine the presence of speech. The additional microphones can be used to perform beamforming to determine whether the user is facing toward a speech source around the user. If the user is facing toward the speech source around the user, the user is most likely having a physical conversation, and no notification is sent. The additional microphones can also be used in beamforming to detect whether the user's head is turned toward a speech source around the user. If the head is turned toward a speech source around the user, the user is most likely having a physical conversation, and no notification is sent. The additional microphones can also be used to estimate whether the headset user is responding to a question from someone around them. The surrounding microphones detect speech around the user, and if the primary microphone detects speech by the user that is estimated to be a response or contribution to a physical conversation in the flow, no notification is sent.
[0015] Introducing a noise cancellation algorithm in combination with one or more additional microphones has been found to improve the efficiency of the VAD algorithm, helping to distinguish between background noise and speech. Furthermore, the use of additional microphones significantly improves the ability to determine when a user is physically speaking with others around them, which can significantly reduce or even eliminate unintentional mute notifications.
[0016] Below, preferred embodiments and features are described, including various ways of determining the described "additional conditions."
[0017] In a preferred embodiment, the method includes determining whether the determined speech appears to originate from a speech source surrounding the user, and providing a mute status notification to the user only if the determined speech does not appear to originate from a speech source surrounding the user. In particular, the method may include processing output signals from the multiple microphones to distinguish between speech by the user and speech from around the user. In particular, the method may include processing output signals from the multiple microphones to provide a beamforming sensitivity pattern to distinguish between speech by the user and speech from around the user.
[0018] In some embodiments, a method includes determining whether a user appears to be engaged in physical conversation and providing a mute state notification to the user only if the user does not appear to be engaging in physical conversation. In particular, the method may include running a first VAD algorithm on an output signal from a microphone, such as a mouth microphone in a headset, that captures the user's speech, and running a second voice activity detection algorithm on an output signal from at least one additional microphone to determine speech from another source. In particular, the method may include determining a time between speech by the user and speech from the another source to determine whether the user appears to be engaging in physical conversation.
[0019] The method may include running a VAD algorithm on a signal indicative of speech of at least one other participant in the call to detect speech by the at least one other participant in the call. In particular, the method may include providing a mute status notification to the user only if the user is speaking and no speech by the at least one other participant in the call is detected. Thus, an intelligent mute status notification may be provided to the user only when it is most likely that the user intends to speak despite the microphone being muted, i.e., when other participants in the call are quiet and the user is speaking.
[0020] The method may include processing the output signal from the main microphone, e.g., a mouth microphone of a headset, and the output signal from the additional microphone to perform a noise cancellation algorithm to suppress ambient noise, which may help improve the performance of the VAD algorithm.
[0021] If desired, the user can configure a frequency meter for mute state notifications, so that if the notifications are still found to be intrusive, the user can reduce the frequency of notifications, thereby experiencing even less disruption.
[0022] A VAD algorithm is understood to detect the presence of speech in a signal. Preferably, features are extracted from the signal in the time or frequency domain and used in classification rules to determine whether speech is present. A microphone, e.g., a headset microphone, provides a real-time signal to the VAD algorithm while remaining muted. Implementation of a VAD algorithm will be known to those skilled in the art.
[0023] In some embodiments, the method includes, by a first processor, performing noise cancellation on signals from the primary microphone and the additional microphone to suppress ambient noise, processing output signals from the microphone system based on a VAD algorithm, determining whether speech is present, and determining whether additional conditions are met, the steps being performed by the first processor, such as a processor in a headset including the primary microphone, the additional microphone, and a speaker. The step of providing a mute status notification is performed by a second processor, such as a processor in a computer device or computer system facilitating the call. In such embodiments, the described steps are preferably used only to determine whether to mute or transmit audio from the primary microphone to the second processor facilitating the call, i.e., to determine whether audio from the primary microphone is transmitted to the second processor facilitating the call if speech is present and the additional conditions are determined to be met. In this way, even when a conventional call system is used, an unintended mute status notification is avoided because a correct mute status notification would not be triggered unless the primary microphone is muted and the user likely intends to speak on the call.
[0024] In some implementations of noise cancellation, the method includes executing a noise cancellation algorithm, including a VAD algorithm that provides an output indicating the presence of speech for the output signal from the primary microphone and the output signal from the additional microphone, and generating a noise-canceled version of the output signal from the primary microphone based on the output indicating the presence of speech. In particular, the noise cancellation algorithm may include applying the output indicating the presence of speech to a noise estimator that estimates noise during periods of speech absence in the output signal from the primary microphone. In particular, the noise cancellation algorithm may include multiplying a frequency-domain representation of the signal from the primary microphone using a set of frequency bins by a gain vector, the gain vector being generated with a lower gain value for frequency bins that do not contain speech and preferably a higher gain value for frequency bins that do contain speech. In particular, the noise cancellation algorithm may include generating the gain vector in response to input from the noise estimator, whereby the gain vector is preferably generated based on noise estimates from the noise estimator.
[0025] It has been found that using a VAD algorithm in a noise cancellation algorithm improves noise estimation because it can be based only on periods of no speech, thereby providing better noise suppression in the signal from the primary microphone, which in turn improves the VAD algorithm implemented to determine mute status notification.
[0026] Another noise cancellation algorithm is based on applying an adaptive noise cancellation algorithm that includes an adaptive filter to generate a noise-canceled version of the output signal from the main microphone. In particular, the adaptive filter can be realized by a least mean squares algorithm or a normalized least mean squares algorithm, as known to those skilled in the art.
[0027] In a second aspect, the present invention provides an apparatus configured for two-way audio communication, such as wirelessly, including a microphone system comprising a primary microphone and an additional microphone, and a processor system configured to perform all steps of the method of the first aspect, or at least the steps of the method of the first aspect other than providing a mute status notification. In particular, the processor system may be configured to determine to mute the primary microphone in response to the additional condition, such that audio output from the primary microphone is provided only when it is determined that a user likely intends to speak in the call. Thus, the apparatus preferably determines to mute the primary microphone to prevent any audio from being transmitted to the processor system facilitating the call, unless it is determined that a user likely intends to speak in the call.
[0028] In particular, the device may be a headset, for example a headset with a processor system forming an integral part thereof.
[0029] The device preferably includes a speaker to allow two-way audio communication. The device may be a standalone device with a microphone system and speaker in one unit, for example, connected wired (e.g., USB) or wirelessly (e.g., Bluetooth) to a computer, smartphone, or the like.
[0030] In particular, the device may include a headset system configured for two-way audio communication, such as wirelessly, the headset system comprising: - a headset configured to be worn by a user, the headset including a microphone system including a mouth microphone, an additional microphone positioned separately from the mouth microphone, and at least one ear cup including a speaker; - A mute enable feature that users can enable during a call to mute the audio from a muted mouth microphone; a processor system configured to perform the method according to the first aspect, or at least the steps other than the step of sending a mute notification, to determine whether it is appropriate to notify a user that the mouth microphone is muted when the user is speaking, or to determine whether the mouth microphone should be muted when the mouth microphone is muted and the user is speaking; In particular, the processor system of the device may preferably be configured to determine whether the user appears to intend to speak and adaptively transmit audio from the mouth microphone only if it is determined that the user appears to intend to speak, in order to prevent any mute state notifications from being sent by an entity such as a processor system facilitating a call.
[0031] In particular, the microphone system may include two or more additional microphones positioned separately from the mouth microphone. For example, the mouth microphone may be implemented as multiple separate microphones to enable beamforming to suppress ambient sounds captured by the mouth microphone. For example, one or several additional microphones may be placed in one or both ear cups of a headset to capture ambient sounds, e.g., for active noise cancellation of sounds reaching the user's ears. For example, the array of additional microphones may be configured to capture speech from a direction limited to the user only and / or to perform beamforming, e.g., to determine the direction from which speech is coming, to enable determination of whether speech appears to be intended for the user as part of a conversation with the user or whether such speech is considered speech not intended for the user.
[0032] In particular, the processor system is configured to provide the notification to the user via a speaker as an audio notification, for example a voice message.
[0033] In particular, the mute function may be implemented as a user-operable knob, push button, contact, or other means located on a portion of the headset.
[0034] The processor system can be a processor known from existing devices such as headsets. Thus, the present invention is suitable for simple implementation in devices with processors with special capabilities, such as for running VAD algorithms. Thus, while it is possible to implement the necessary processing in a small headset, if desired, the processing system can be implemented on a computer or smartphone, or on a dedicated device separate from the headset.
[0035] In a third aspect, the present invention provides a communication system, the communication system comprising: at least one device according to the first aspect; - a communications device configured to provide two-way communication over a communications channel and optionally provide two-way audio over a digital wireless system, such as DECT, Bluetooth or other similar short-range wireless system, to at least one device according to the first aspect; Includes:
[0036] In particular, the communication device may include a computer or a mobile phone such as a smartphone. The communication channel may be a cellular network, e.g., 2G, 3G, 4G, 5G, or the like, the Internet, or a dedicated communication channel, wired or wireless, etc. The connection between the communication device and the communication channel may be a wired connection or a wireless connection, and the connection may include, for example, a wi-fi connection.
[0037] In particular, the communication system may be, for example, a teleconference system or the like.
[0038] In a fourth aspect, the present invention provides use of the method according to the first aspect for conducting one or more of a telephone call, an online call, or a conference call.
[0039] In a fifth aspect, the present invention provides use of an apparatus according to the second aspect for conducting one or more of a telephone call, an online call, or a conference call.
[0040] In a sixth aspect, the present invention provides use of a system according to the third aspect for conducting one or more of a telephone call, an online call, or a conference call.
[0041] In a seventh aspect, the present invention provides program code, which when executed on one or two separate processors, is configured to cause the execution of a method according to the first aspect. In particular, the program code can be stored in memory on a chip, or on one or more tangible storage media, or can be available on the internet in a version for download. The program code can be in general code format or in a processor-specific format.
[0042] It is understood that the same advantages and embodiments described for the first aspect apply equally to the further described aspects, and further, it is understood that the described embodiments can be mixed in any way between all the described aspects. [Brief explanation of the drawings]
[0043] The present invention will now be described in more detail with reference to the accompanying drawings.
[0044] [Figure 1] This illustrates a situation where a headset user is participating in an online call, but is also physically present in the room with other people who are speaking to the headset user during the call. [Figure 2] 1 illustrates steps of a method embodiment. [Figure 3]1 shows a block diagram with elements of an embodiment; [Figure 4] 1 illustrates an embodiment of a headset system. [Figure 5] 1 shows a block diagram of elements of an embodiment including noise cancellation provided at both the main microphone (mouth microphone) and the additional microphones prior to providing signals from those microphones to the VAD algorithm. [Figure 6] 1 illustrates an embodiment of a headset system that includes an additional microphone located in the ear cup and a processor that decides to transmit audio output from the primary microphone (mouth microphone) only if it is determined that the user appears to intend to speak in an ongoing call. [Figure 7] 1 shows a block diagram of an example noise cancellation algorithm that generates a noise-canceled version of the audio signal from the main microphone based on audio inputs from the main microphone and additional microphones. [Figure 8] 1 shows a block diagram of another example of a noise cancellation algorithm based on adaptive noise cancellation.
[0045] The figures illustrate particular ways of practicing the invention and should not be construed as limiting other possible embodiments that fall within the scope of the appended set of claims. DETAILED DESCRIPTION OF THE INVENTION
[0046] 1 illustrates the basic situation behind the present invention: a user U is present in a physical room RM with another person P, e.g., a colleague. The user U is engaged in a call CL, e.g., an online conference using a computer or the like, with the other call participant CL_P. The user U is wearing a headset for two-way communication with the call participant CL_P. If the user U mutes the headset microphone for some reason and noise or speech is captured by the headset's mouth microphone, a mute status notification, either a visual message on the display or an audio message via the headset's speaker, is provided to the user. However, such a notification may be unintended and disruptive to the user, e.g., if the captured audio is speech by the person P in the room RM and / or speech by the user U in a conversation with the person P in the room RM.
[0047] This problem is solved in accordance with the present invention by using a voice activity detection (VAD) algorithm and additional criteria to determine whether a mute status notification should be provided to the user U. This makes it possible to eliminate notifications that are unintended and may disturb the user U rather than serving as an aid.
[0048] 2 illustrates steps of a method embodiment, i.e., a method for notifying a user of a muted microphone system when the user speaks while the microphone system is muted during a call with one or more other participants. The method includes processing an output signal from a primary microphone, e.g., a mouth microphone of a headset, and an output signal from an additional microphone to execute an environmental noise cancellation algorithm (ENC) to suppress ambient noise from the user's environment. The method further includes processing the microphone system, at least the output signal from the primary microphone, and optionally the output signals from both the primary microphone and the additional microphone, based on a VAD algorithm using a processor system while the microphone system is muted (VAD). Next, determining whether speech is present based on the output of the VAD algorithm (S_D). Furthermore, determining whether additional conditions are met (D_AC), apart from whether speech may be detected, and finally providing a mute status notification to the user only if it is determined that speech is present and the additional conditions are met (P_MSN).
[0049] In some embodiments, steps ENC, VAD, S_D, and D_AC are performed by a first processor in a first device, e.g., a headset, and step P_MSN is performed by a second processor in a second device, e.g., a computer conducting a call with a remote party. In some embodiments, all five steps described are performed by a processor in a single device.
[0050] The additional conditions may be based on one or more additional VAD algorithms configured to operate on additional microphones to determine whether speech is present in the user's environment and / or on audio input from the call to determine whether other participants are speaking, which may be useful in providing important information to determine the actual situation in which the user finds themselves and whether it is appropriate to provide a mute status notification.
[0051] Noise cancellation algorithms (often referred to as ENC, ANC, or the like) are used to improve the performance of one or more VAD algorithms.
[0052] The described method may be implemented in, for example, a headset to utilize a method for intelligently providing mute status notification.
[0053] A block diagram illustrating a portion of an embodiment of a headset is shown in Figure 3. A decision algorithm D_A determines whether a mute state notification MT_N should be sent to the user if certain conditions are met and the user's mouth microphone MM is in a mute state MT, i.e., blocking the user's audio during an ongoing call.
[0054] The first VAD algorithm VAD1 operates on the signal from the mouth microphone MM of the headset and determines the first input to the decision algorithm D_A, i.e., whether speech is present. The second VAD algorithm VAD2 operates on input from one or more microphones configured to capture sound from the environment around the user, for example one or several microphones located on the exterior of the headset, and provides the presence or absence of speech in that environment to the decision algorithm D_A. Finally, the third VAD algorithm VAD3 operates on the audio input from the call CS and is responsible for determining whether other participants in the call are talking or quiet.
[0055] Therefore, the decision algorithm D_A has two inputs, from VAD2 and VAD3, in addition to the input from VAD1, which can assume that the user is speaking. In particular, the input from VAD2 can be used to determine whether someone in the environment is speaking while the user is speaking, which most likely means that the user is conversing with someone present in the environment and may not intend to speak to the participants in the call; therefore, the mute status notification MT_N should be avoided in such a case. Furthermore, if the user is detected to be speaking, the call audio CS indicates that no other participants are speaking, which means that the user likely wants to speak in the call; therefore, it is appropriate to provide the mute status notification MT_N.
[0056] 4 illustrates an embodiment of a headset system comprising a headset HS worn by a user during a call, the headset having a primary microphone in the form of a mouth microphone MM for capturing the user's voice and two ear cups each equipped with a speaker for providing audio from a call CL to the user. The mouth microphone MM and speaker of the headset HS are connected to a processor P, which may be integrated into one or both ear cups of the headset HS, for example. The processor P handles two-way audio communication associated with the call CL, such as in a wireless manner. The headset HS has a mute enable function MT that the user can enable during the call CL to mute the audio from the mouth microphone MM in a mute state MT. The mute state MT is provided as an input to the processor P, which provides a mute state notification MT_N to the user when appropriate, in accordance with the method described above, i.e., only when the VAD algorithm detects that the user is speaking with the mouth microphone MM in a mute state MT.
[0057] It should be understood that the illustrated embodiment of the headset system is configured for two-way voice communication CL over a wired or wireless communication channel to a communication device that is responsible for providing the call connection.
[0058] In some headset system embodiments, at least a portion of the mute function for muting audio from the primary microphone or mouth microphone is implemented on a processor forming part of the headset system. Thus, in such embodiments, the headset system simply mutes the primary microphone itself if it determines that the user appears to believe the primary microphone should be muted. Thus, such embodiments are compatible with existing communications devices or computer programs responsible for providing call connections over a communications channel, since such devices or programs are prompted to send a mute notification only when the headset passes audio that is likely to be the user's speech for the call, and as a result, the device or program's mute notification functions as intended, i.e., with improved quality compared to that with standard headset systems. It should be understood, however, that the processing and mute notification decisions may, in other embodiments, be performed entirely by the device or program facilitating the call.
[0059] The following four sub-aspects 1)-4) have been found to improve the performance of the mute status notification method and apparatus and are therefore considered preferred embodiments.
[0060] 1) Context Awareness with Beamforming. This utilizes additional microphones installed in the headset to function as a microphone array. Beamforming technology is used to directionally identify a person, such as a colleague, who is speaking in the user's environment. If the person is detected within a certain reception angle, the method can be configured to identify the likely context as a conversation with that person, and as a result, it is determined that a mute status notification should not be provided. Alternatively, or in addition, the beamforming configuration can be used to detect whether the user is paying attention to the person. This is done by using beamforming to detect whether the user is turning their head toward the person who is speaking. When the person begins to speak, the headset detects the person at a certain angle. If the user turns their head toward the person, the headset will detect the person at a different angle, and as a result, it may be determined that a conversation is the likely context and that a mute status notification should not be provided.
[0061] 2) Noise cancellation algorithms that optimize VAD performance. For example, an environmental noise cancellation (ENC) algorithm can use inputs from a main microphone (e.g., a mouth microphone) and one or more separate microphones to remove ambient noise. By combining the two approaches, the VAD algorithm is less affected by ambient noise, thereby reducing the risk that ambient sounds will falsely activate a mute notification.
[0062] 3) Conversational Context Awareness. A primary microphone (e.g., a mouth microphone) can be used to capture the user's speech, and one or more auxiliary microphones can be used to capture speech around the user. For each input at the two microphones, a VAD algorithm running separately detects whether speech is present and notifies the headset when the user is speaking and when someone is speaking around the user. A model can be used to assess the likelihood that speech captured by the two microphones is part of the same conversation. This assessment can be used to determine whether a mute status notification should be provided.
[0063] 4) Call Activity Context Recognition. When a user is speaking, two separately running VAD algorithms are used: one VAD algorithm detects speech in the signal from the primary microphone (e.g., the mouth microphone); the other VAD algorithm processes the audio input from the call to detect speech and determine the call activity, i.e., speech activity in the call. The presence of speech in the call activity is used to assess the likelihood that the user is unintentionally speaking into a muted microphone. If no speech is detected in the call activity and the user is speaking into a muted microphone, it is estimated that the call participants are likely waiting for the user to contribute, and a mute status notification is provided. If speech is detected in the call activity and the user is speaking into a muted microphone, it is estimated that the call participants are less likely waiting for the user to contribute, and no mute status notification is provided in such cases.
[0064] 5 shows a block diagram illustrating part of an embodiment of a headset with a mouth microphone MM as a main microphone and an additional microphone M2. A decision algorithm D_A decides whether to mute the audio from the mouth microphone MM or pass the audio from the mouth microphone MM to the audio output A_O depending on whether certain conditions are met.
[0065] The audio output from the mouth microphone MM and the additional microphone M2 are both processed by a noise cancellation algorithm NC to cancel noise that may be present in the audio output from the mouth microphone MM, and the noise-suppressed audio signal from the mouth microphone MM is provided as input to a VAD algorithm VAD1. The audio output from the additional microphone M2 is processed by another VAD algorithm VAD2. It will be appreciated that separate noise estimation algorithms can instead be provided for the outputs from the two microphones MM and M2, if desired.
[0066] Each of the VAD algorithms VAD1 and VAD2 provides a result that is provided as an input to a decision algorithm D_A, i.e., an algorithm that determines whether speech is present at each of the two microphones MM and M2. These inputs can be used, among other things, to determine whether the user is talking to someone in the environment, i.e., appears to be having a physical conversation with another person. In such a case, the decision algorithm D_A makes the decision to mute the audio from the mouth microphone and provide the speech from the mouth microphone at audio output A_O if it is detected that the user is speaking based on VAD1 and VAD2 indicates that there is no further speech in the environment for a certain period of time.
[0067] FIG. 6 illustrates a variation of the headset system (enclosed by the dashed line) of FIG. 4 . In FIG. 6 , the headset HS includes a primary microphone, here shown as a mouth microphone MM, and an additional microphone AM, here located in an ear cup of the headset HS, for capturing ambient sounds. A processor system P1, implemented, for example, integrally with one of the ear cups of the headset HS, is configured to process output signals from the mouth microphone MM and the additional microphone AM to execute a noise cancellation algorithm to suppress ambient noise. Furthermore, the processor system P1 is configured to process the output from the mouth microphone MM based on a VAD algorithm, and optionally, the output from the additional microphone is also processed based on another VAD algorithm, for example, as in FIG. 5 . Furthermore, the processor system P1 is configured to determine whether speech is present based on the output of the VAD executed on the output from the mouth microphone MM, and further determine whether an additional condition is satisfied. The processor system P1 is configured to generate an audio output A_O from the mouth microphone MM only if it is determined that the mouth microphone MM is capturing speech and the additional condition is satisfied. In particular, the additional condition may be that it is determined that the user is likely speaking and that the user is not engaged in a physical conversation with people around them. Specifically, the determination of the additional condition may be based on processing of audio captured by the additional microphone AM.
[0068] Another processor system P2 facilitates the call, thereby providing two-way audio connectivity to the call participant CL_P. The processor system P2 may comprise a personal computer, laptop, tablet, smartphone, or dedicated device, and is responsible for processing the audio output A_O from the headset system and generating an audio input A_I to the headset system using audio from the distant call participant CL_P.
[0069] In this way, existing general-purpose call or online communication programs can be used with the headset and still obtain the functionality of the more intelligent mute notification MT_N, because the separate processor system P2 provides the mute notification MT_N in the conventional manner known from existing call systems, for example, when the audio level of the audio output A_O exceeds a certain level in a muted state. The notification MT_N can be provided, for example, as a visual and / or audio notification. However, since the processor system P1 of the headset system is responsible for providing intelligent muting of the mouth microphone MM, it is ensured that the audio output A_O to the separate processor system P2 is only provided when the headset system determines that the user appears to intend to speak in an ongoing call, thereby eliminating the annoying mute state notification MT_N even in existing call systems.
[0070] FIG. 7 shows an example of a noise cancellation algorithm that processes an audio signal A_MM from a primary microphone and an audio signal A_M2 from an additional microphone to generate a noise-canceled audio signal A_MM_NC from the primary microphone. Essentially, the algorithm operates on frequency-domain representations X and X2 of the audio input signals A_MM and A_M2, respectively. A gain vector G is multiplied by the frequency representation of the primary microphone audio signal X. The gain vector G is generated such that a lower gain is assigned to frequency bins in the frequency representation of the primary microphone signal X that do not contain speech. The output Y resulting from the multiplication of X and G is then converted to a time signal A_MM_NC that represents a noise-canceled version of the original audio signal A_MM from the primary microphone.
[0071] 7 shows an initial short-term analysis STA performed on each speech signal A_MM and A_M2, based on which the two speech signals A_MM and A_M2 are converted into their respective frequency domain representations X and X2. X is applied to a noise estimator NE which estimates the noise N to generate a gain vector G which amplifies frequency bins containing speech and attenuates frequency bins not containing speech, and finally the gain estimator GE generates a gain vector G based on the estimated noise N and X. The noise estimator NE receives an input V from a voice activity detector VAD which operates on both X and X2 as inputs, the input V indicating to the noise estimator NE when speech is or is not present, and the noise estimator NE updates its noise estimate N during periods when speech is not present.
[0072] 8 shows a block diagram of another example of a noise cancellation algorithm based on simple adaptive noise cancellation. This algorithm is based on the assumption that the audio signal x from the main microphone contains the intended speech and noise, and the audio signal x2 from the additional microphone contains the same noise, but this assumption may not be completely valid in practice because the two microphones are located at different locations.
[0073] The objective of adaptive noise cancellation is to minimize the output power z. This is achieved by using the output signal as the error signal e in the adaptive filter AF. It can be proven that the minimum output power is achieved when y equals the noise, which means that the output signal z equals the desired signal x.
[0074] Several algorithms can be used as adaptive filters AF, such as the normalized least mean squares (NLMS) algorithm, which is based on the least mean squares (LMS) algorithm, where gradient descent is used to adjust the filter coefficients to minimize the error e. NLMS normalizes the power of the input and uses a time-varying step size for faster convergence.
[0075] It is understood that the described noise cancellation examples only serve to illustrate that noise cancellation of the audio signal from the primary microphone can be achieved in various ways, and therefore the effect of improving the reliability of the VAD performed on the noise-canceled primary microphone signal can be achieved through various implementations.
[0076] In the following, additional embodiments E1 to E15 are defined.
[0077] E1. A method for notifying a user of a muted state of a microphone system when the user speaks while the microphone system is muted during a call with one or more other participants, comprising: - processing an output signal from the microphone system with a processor system based on a voice activity detection algorithm (VAD) while the microphone system is muted; - determining whether speech is present based on the output of a voice activity detection algorithm (S_D); - determining whether additional conditions are met (D_AC); - Providing a mute status notification to the user only when speech is present and additional conditions are determined to be met (P_MSN); A method comprising:
[0078] E2. The method of E1, including determining whether the determined speech appears to originate from a speech source surrounding the user, and providing a mute status notification to the user only if the determined speech does not appear to originate from a speech source surrounding the user.
[0079] E3. The method of E2, including processing output signals from a plurality of microphones to distinguish between speech by the user and speech from around the user.
[0080] E4. The method of E3, including processing output signals from a plurality of microphones to provide a beamforming sensitivity pattern to distinguish between speech by a user and speech from around the user.
[0081] E5. A method according to any one of E1 to E4, including determining whether the user appears to be having a physical conversation and providing a mute status notification to the user only if the user does not appear to be having a physical conversation.
[0082] E6. The method of E5, including running a first voice activity detection algorithm on an output signal from a microphone, such as a mouth microphone, that captures a user's speech, and running a second voice activity detection algorithm on an output signal from at least one additional microphone to determine speech from another source.
[0083] E7. The method of E5 or E6, including determining the time between speech by the user and speech from another source to determine whether the user appears to be engaged in physical conversation.
[0084] E8. The method of any one of E1-E7, comprising running a voice activity detection algorithm on a signal indicative of speech from at least one other participant in the call to detect speech by at least one other participant in the call.
[0085] E9. The method of E8, including providing a mute status notification to the user only when the user is speaking and speech by at least one other participant in the call is not detected.
[0086] E10. A method according to any one of E1 to E9, comprising processing an output signal from a main microphone, e.g. a mouth microphone of a headset, and an output signal from an additional microphone to perform a noise cancellation algorithm (ENC) to suppress ambient noise.
[0087] E11. An apparatus comprising a microphone system and a processor system (P), the apparatus being configured to perform the method of any one of E1 to E10.
[0088] E12. The device according to E11, comprising a headset system configured for two-way voice communication, such as wirelessly, the headset system comprising: - a headset (HS) adapted to be worn by a user, the headset (HS) including a microphone system comprising at least a mouth microphone (MM) and at least one ear cup with a speaker; - Mute enable function (MT) that can be enabled by the user during a call to mute the audio from the muted mouth microphone (MM); a processor system (P) configured to perform a method according to any one of E1 to E10 to determine whether it is appropriate to notify a user of a mute state notification when the mouth microphone (MM) is muted and the user is speaking; 1. An apparatus comprising:
[0089] E13. The apparatus according to E12, wherein the microphone system comprises at least one additional microphone (M2) positioned separately from the mouth microphone (MM).
[0090] E14. The apparatus according to E12 or E13, wherein the processor system (P) is configured to provide the notification to the user as an audio notification via a speaker.
[0091] E15. Use of the method of any one of E1 to E10 to conduct one or more of a telephone call, an online call, or a conference call.
[0092] In summary, the present invention provides a method and apparatus, e.g., a headset, for notifying a user of the muted state of a primary microphone during a call when the user speaks with the primary microphone muted. The method includes executing a noise cancellation algorithm (ENC) on an output signal from the primary microphone and an output signal from an additional microphone capturing audio around the user to suppress ambient noise at the user's location. Furthermore, the output signal from the primary microphone when muted is processed based on a voice activity detection (VAD) algorithm using a processor system. The VAD algorithm is used to determine whether speech is present, and then it is determined whether additional conditions are met. Finally, a mute status notification is provided to the user only if speech is present and the additional conditions are met. This method is well suited for headsets, for example, where various noises at the mouth microphone may otherwise cause unintended and distracting mute status notifications. Using the VAD algorithm, it is possible to ensure that notifications are triggered only by speech, and additional conditions can be used to eliminate or at least reduce intrusive notifications by intelligently providing mute status notifications based on, for example, speech activity of other participants on the call and based on speech around the user.
[0093] Although the present invention has been described with reference to specific embodiments, it should not be construed as being limited to the examples provided. The scope of the present invention should be interpreted in light of the appended set of claims. In connection with the claims, the terms "comprising" or "includes" do not exclude other possible elements or steps. Furthermore, references such as "a," "an," etc. should not be construed as excluding a plurality. Furthermore, the use of reference signs in the claims to elements shown in the figures should not be construed as limiting the scope of the present invention. Furthermore, individual features recited in different claims may in some cases be advantageously combined, and reference to such features in different claims does not exclude that combination of features is possible or advantageous.
Claims
1. 1. A method of informing a user of a muted state of a primary microphone configured to capture speech of the user during a call with one or more other participants when the user speaks while the primary microphone of a microphone system is muted, comprising: 1) processing the output signal from the main microphone and the output signal from an additional microphone positioned to capture sounds around the user to perform a noise cancellation algorithm to suppress ambient noise; - 2) processing the output signal from the primary microphone with the primary microphone muted based on a voice activity detection algorithm using a processor system; 3) determining whether speech is present based on the output of the voice activity detection algorithm; 4) determining whether additional conditions are met; and 5) providing a mute status notification to the user only if speech is present and the additional condition is determined to be met; and Including, The method of determining the additional condition includes determining whether the determined speech appears to originate from a speech source surrounding the user, and providing the mute status notification to the user only if the determined speech does not appear to originate from a speech source surrounding the user.
2. The method of claim 1 , comprising processing output signals from a plurality of microphones to distinguish between speech by the user and speech from around the user.
3. The method of claim 2 , comprising processing the output signals from the plurality of microphones to provide beamforming sensitivity patterns capable of distinguishing between speech by the user and speech from around the user.
4. 10. The method of claim 1, comprising: executing a voice activity detection algorithm on a signal indicative of the voice of at least one other participant in the call to detect speech by the at least one other participant in the call.
5. 5. The method of claim 4, comprising providing the mute status notification to the user only if the user is speaking and no speech is detected by at least one of the other participants in the call.
6. 10. The method of claim 1, wherein steps 1) through 4) are performed by a first processor, such as a processor of a headset that includes the primary microphone, the additional microphone, and a speaker, and step 5) is performed by a second processor, such as a processor of a computer device or computer system that facilitates the call.
7. 2. The method of claim 1, further comprising, after steps 1) to 4), making a decision to mute audio from the primary microphone when it is determined that speech is present and the additional condition is met, to prevent transmission of the mute state notification.
8. 2. The method of claim 1, comprising: executing a noise cancellation algorithm including a voice activity detection algorithm that provides an output indicative of the presence of speech on the output signals from the primary microphone and the additional microphone; and generating a noise-canceled version of the output signal from the primary microphone based on the output indicative of the presence of speech.
9. 9. The method of claim 8, comprising applying the output indicating the presence of speech to a noise estimator that estimates noise during periods when no speech is present in the output signal from the primary microphone.
10. The method of claim 1 , comprising applying an adaptive noise cancellation algorithm including an adaptive filter to generate a noise-canceled version of the output signal from the main microphone.
11. An apparatus including a microphone system comprising a main microphone and an additional microphone, and a processor system configured to perform at least steps 1) to 4) of the method of claim 1.
12. 12. The device of claim 11, wherein the processor system is configured to determine to mute the primary microphone in response to the additional condition, such that audio output from the primary microphone is provided only when it is determined that the user appears to intend to speak on the call.
13. a headset system configured for two-way voice communication, such as wireless communication, the headset system comprising: a headset configured to be worn by the user, the headset including a microphone system comprising a mouth microphone, an additional microphone positioned separately from the mouth microphone, and at least one ear cup comprising a speaker; a mute enable function that the user can enable during the call to mute the audio from the mouth microphone that is muted; a processor system configured to perform at least steps 1) to 4) of the method of claim 1 to determine whether it is appropriate to notify the user that the mouth microphone is in the muted state when the user is speaking with the mouth microphone in the muted state, or to determine whether the mouth microphone should be muted when the user is speaking with the mouth microphone in the muted state; 13. The apparatus of claim 12, comprising:
14. 14. The device of claim 13, wherein the processor system is configured to determine whether the user appears to intend to speak and to adaptively transmit audio from the mouth microphone only if it is determined that the user appears to intend to speak to prevent any mute state notifications from being sent by an entity facilitating the call.
15. 12. Use of the device of claim 11 to conduct one or more of a telephone call, an online call, or a conference call.
Citation Information
Patent Citations
Method for detecting user voice activity in a communication assembly, said communication assembly
JP2020506634A
Microphone mute / unmute notification
US20150156598A1
System and method for providing intelligent and automatic mute notification
US20150195411A1