Wearable electronic device for transmitting masking signals
By integrating electroacoustic input transducers, speakers and processors in wearable electronic devices, dynamically adjusting the volume of the masked signal in response to the voice activation signal, the problem of difficulty in effectively masking voice interference in the prior art is solved, and more efficient noise masking and wearer comfort is achieved.
Patent Information
- Application Number
- CN202011064664.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-04
- Filing Date
- 2020-09-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-09-30
AI Technical Summary
Existing headphones and in-ear headphones are not effective in reducing human voice interference in the surrounding environment, especially under active noise reduction technology, where voice-activated noise remains difficult to effectively mask, resulting in the wearer still being disturbed while performing cognitive tasks.
A wearable electronic device is designed, including an electroacoustic input transducer, a speaker and a processor. The processor generates a voice activation signal by detecting voice activation in the microphone signal, and adjusts the volume of the masking signal according to the signal, providing a masking signal with a larger first volume when voice activation is detected, otherwise providing a smaller second volume or a stop masking signal.
By dynamically adjusting the volume of the masking signal, it can effectively mask voice activation noise, reduce interference to the wearer's ears, reduce auditory fatigue, and improve the quietness of the wearer's working environment.
Smart Images

Figure CN112616105B_ABST
Abstract
Description
Technical Field
[0001] Wearable electronic devices such as headphones or in-ear headphones include a pair of small speakers, which are located in different ways in the earpieces worn by the wearer (user of the wearable electronic device), depending on the configuration of the headphones or in-ear headphones. In-ear headphones are usually placed at least partially in the ear canal of the wearer, while headphones are usually worn by a headband or neckband, and the earpieces are placed on or above the wearer's ears. In contrast to traditional speakers, headphones or in-ear headphones allow the wearer to listen to the audio source privately, while traditional speakers emit sound outdoors for anyone nearby to listen. Headphones or in-ear headphones may be connected to an audio source to play audio. In addition, headphones can be used to establish a private quiet space, such as by one or both of passive or active noise reduction, to reduce the tension and fatigue generated by the wearer due to the sound in the surrounding environment. In an open office environment where other people are talking (such as, talking loudly), wearable electronic devices (such as headphones) can be used to obtain a quiet working environment. However, it has been found that both passive and active noise reduction are insufficient to reduce the distracting characteristics of human speech in the surrounding environment. This distraction is most often caused by nearby people talking, for example, while the user is performing a cognitive task, although other sounds may also distract the user.
[0002] In particular, this can be a problem for active noise cancellation, which excels at reducing tonal or low-frequency noise (e.g. from machinery) but is less effective at reducing voice-activated noise. Active noise cancellation depends on capturing a microphone signal, e.g. in a feedback, feedforward or hybrid fashion, and sending a signal through a speaker to cancel out the ambient sound (noise) signal from the surrounding environment.
[0003] In contrast, conventionally, in a telecommunications environment, a headset enables communication with a remote party, for example via a phone (which may be a so-called softphone or another type of application running on an electronic device). The headset may use wireless communication, for example, according to Bluetooth or DECT compatible standards. However, the headset relies on capturing the wearer's own voice to transmit the voice signal to the remote party. Background Art
[0004] Headphones or earphones with active noise reduction or active noise cancellation (sometimes abbreviated as ANC or ANR) can help provide the wearer with a quieter private work environment, but such devices are limited in that they do not reduce the speech of nearby people to a level that is inaudible and unintelligible. Therefore, some level of distraction still exists.
[0005] It has been shown that playing instrumental music to a person can reduce distractions caused by people speaking near that person to some extent. However, if the intensity of the distracting sounds varies over the course of a day, trying to listen to music at a fixed volume to mask distracting speech activations may not be ideal. High levels of instrumental music may mask all distracting sounds, but listening to music at this level for long periods of time may cause auditory fatigue. On the other hand, soft music levels may not adequately mask distracting sounds.
[0006] US8,964,997 (licensed to Bose Corporation) discloses a masking module that automatically adjusts the audio level to reduce or eliminate the interference or other effects of residual ambient noise in the earpiece on the user. The masking module uses an audio signal presented through the headphone to mask the ambient noise. The masking module performs gain control and / or level compression based on the noise level, so that the user is not easily aware of the surrounding noise. In particular, the masking module adjusts the level of the masking signal so that it is only as large as required to mask the residual noise. The value of the masking signal is determined experimentally to provide sufficient masking for interfering speech. Therefore, the masking module uses the masking signal to provide additional isolation over the active or passive attenuation provided by the headphone.
[0007] US2015 / 0348530 (licensed to Plantronics) discloses a system for masking interfering sounds in headphones. The noise masking signal essentially replaces meaningful but unwanted sounds (such as human speech) with useless and therefore less interfering noise (i.e., so-called "comfort noise"). When the surrounding noise subsides (e.g., when the interfering sound ends), the digital signal processor automatically gradually decays the noise masking signal back to silence. The digital signal processor uses dynamic or adaptive noise masking, so that as the interfering sound increases (e.g., a speaker approaches the headphones), the digital signal processor increases the noise masking signal with the amplitude and frequency response of the interfering signal. It is emphasized that the embodiments are intended to reduce the intelligibility of ambient speech while having no deleterious effect on the headphone audio speech intelligibility.
[0008] However, there remains the problem that the headphone wearer may suffer from unpleasant listening fatigue because the speaker will emit a masking signal whenever an interfering sound is detected. Summary of the invention
[0009] Therefore, there is a need for a wearable device that masks distracting noises but at the same time minimizes listening fatigue. Provided are:
[0010] A wearable electronic device, comprising:
[0011] an electroacoustic input transducer arranged to pick up an acoustic signal and convert the acoustic signal into a microphone signal;
[0012] Speakers; and
[0013] A processor configured to:
[0014] controlling the volume of the masking signal; and
[0015] providing a masking signal to a loudspeaker;
[0016] The processor is further configured to:
[0017] Based on processing at least the microphone signal, detecting voice activation and generating a voice activation signal concurrently with the microphone signal, the voice activation signal sequentially indicating one or more of: voice activation and voice inactivity; and
[0018] In response to the voice activation signal, a volume of the masking signal is controlled according to providing the masking signal to the speaker at a first volume when the voice activation signal indicates voice activation and providing the masking signal to the speaker at a second volume when the voice activation signal indicates voice inactivity.
[0019] In some aspects, the first volume is greater than the second volume. In some aspects, the first volume is always at a level higher than the second volume. In some aspects, based on the voice activation signal, a masking signal is provided to the speaker when voice activation is currently present. The masking signal acts to actively mask the talk signal, which may leak into one or both ears of the wearer despite some passive suppression by the wearable device. The passive suppression may be caused by the wearable electronic device occupying the wearer's ear canal or arranged on or around the wearer's ear. In response to the voice activation signal, active masking is achieved by controlling the volume of the masking signal. The volume of the masking signal is louder when voice activation is detected than when voice inactivity is detected.
[0020] Thus, by providing the masking signal (at a first volume) to the speaker when the voice activation signal indicates voice activation, the masking effect of the intelligibility of the conversational speech is enhanced or activated. Sometimes, when the voice activation signal indicates that the speech is not active, the volume of the masking signal is reduced (at a second volume) or stopped (corresponding to a second volume that is infinitely less than the first volume). Therefore, when the voice activation signal indicates that the speech is not active, the volume of the masking signal is reduced because the voice-activated masking is not needed to reduce the intelligibility of the conversation near the wearer.
[0021] In some examples, the second volume corresponds to stopping providing the masking signal to the speaker or providing the masking signal at a level considered barely audible to a user with normal hearing. In some examples, the second volume is significantly less than the first volume, for example, 12-50 dB-A lower than the first volume.
[0022] Thus, during a day of use or less, the user is exposed to the masking signal only when the masking signal serves to reduce the intelligibility of speech reaching the ear of the headphone wearer. This in turn reduces auditory fatigue caused by the masking signal emitted by the speaker during a day of use or less. As a result, the wearer is subjected to less sound stress.
[0023] Thus, the wearable device may react to other sounds in the work environment (such as keystrokes on a keyboard) by emitting a masking signal at a first volume sufficient to mask the ambient voice activation, but not masking at all or only emitting a masking signal at a lower second volume, thereby utilizing sounds other than those associated with the conversation, which tend to be less distracting to a person than audible conversation.
[0024] When people are speaking near the wearer (e.g., within a range of up to 8 to 12 meters), the wearable electronic device may send a masking signal to the wearer's ears. The range depends on the threshold sound pressure at which voice activation is detected. Such a threshold sound pressure may be stored or implemented by the processor. The range also depends on the volume of the voice activation, i.e., how loudly one or more people are speaking.
[0025] In some aspects, when the voice activation signal indicates voice activation, the volume of the masking signal is adjusted according to the sound pressure level of the acoustic signal picked up by the electroacoustic input transducer when the voice activation signal indicates voice activation.
[0026] In some examples, when the voice activation signal indicates voice activation, the volume of the masking signal is adjusted based on the sound pressure level of the acoustic signal picked up by the electroacoustic input transducer when the voice activation signal indicates voice activation. For example, the volume of the masking signal is adjusted proportionally to the sound pressure level of the acoustic signal picked up by the electroacoustic input transducer when the voice activation signal indicates voice activation. In some examples, at least when the sound pressure level is below a predetermined upper threshold and / or above a predetermined lower threshold, the volume of the masking signal is adjusted proportionally (e.g., substantially linearly or stepwise) to the sound pressure level of the sound signal. In some aspects, the masking signal is a two-level signal controlled to have a first volume or a second volume. In some aspects, the masking signal is a three-level signal controlled to have a first volume or a second volume or a third volume. The first volume can be a fixed first volume. The second volume can be a fixed second volume, for example corresponding to "off" not being provided to the speaker. The third volume can be higher or lower than the first volume or the second volume. In some aspects, the masking signal is a multi-level signal with more than three volume levels.
[0027] In some aspects, for example, when the voice activation signal indicates voice activation, the volume of the masking signal is adaptively controlled in response to the sound pressure level of the sound signal. In some aspects, when the voice activation signal indicates voice inactivity, the processor or method stops adaptively controlling the volume of the masking signal.
[0028] In some aspects, the processor concurrently:
[0029] - providing a masking signal to a speaker and / or controlling the volume of the masking signal in response to the voice activation signal; and
[0030] - Stop signal processing that enables sound captured by a microphone at the wearable device to be passed to a speaker of the wearable electronic device.
[0031] In some aspects, the processor concurrently:
[0032] - providing a masking signal to a speaker and / or controlling the volume of the masking signal in response to the voice activation signal; and
[0033] - stopping signal processing that causes sound captured by a microphone at the wearable device to be transmitted to a speaker of the wearable electronic device; and
[0034] -Perform active noise cancellation.
[0035] When speech is not detected but noise such as keyboard pressing may be present, the wearable electronic device may stop transmitting masking signals to the wearer's ears. This may be the case in an open office environment. The wearable electronic device may be configured as a headset or a pair of in-ear headphones, for example, and may be used by the wearer of the device to obtain a quiet working environment in which the detected audio speech signals reaching the wearer's ears are masked.
[0036] The processor may be implemented as known in the art and may include a so-called voice activity detector (commonly abbreviated as VAD, voice activity detector, voice endpoint detector), also known as a talk activation detector or talk detector. The voice activation detector is capable of distinguishing periods of voice activity from periods of voice inactivity. Voice activity may be considered to be a state in which the processor can detect the presence of human talk. Voice inactivity may be considered to be a state in which the processor cannot detect the presence of human talk. The processor may perform one or both of time domain processing and frequency domain processing to generate a voice activation signal.
[0037] The voice activation signal may be a binary signal, wherein voice activation and voice inactivation are represented by corresponding binary values. The voice activation signal may be a multi-level voice activation signal representing, for example, one or both of the following: the likelihood of conversation activation occurring; and the level of detected voice activation, such as loudness. In response to the multi-level voice activation signal, the volume of the masking signal may be gradually controlled over more than two levels. In some aspects, the processor is configured to adaptively control the volume of the masking signal in response to the microphone signal. In some aspects, the volume of the masking signal is set according to an estimated required masking volume. The volume of the masking signal may, for example, be set equal to the estimated required masking volume or according to another predetermined relationship. The estimated required masking volume may be a function of one or both of the following: an estimated volume of conversation activation; and an estimated volume of activations other than conversation activation. The estimated required masking volume may be proportional to the estimated volume of conversation activation. The estimated required masking volume may be obtained from experiments, for example, by conducting a hearing test to determine the volume of the masking signal, which is at least sufficient to reduce the distraction of conversation activation to a desired level. The estimated volume of talk activation and / or the estimated volume of activation other than talk activation can be determined based on the processing of the microphone signals. In some aspects, the processing can include processing a beamforming signal (the signal is obtained by processing multiple microphone signals from corresponding multiple microphones).
[0038] The voice activation signal is concurrent with the microphone signal, but the signal processing for detecting voice activation takes some time to perform, so the voice activation signal encounters a delay in detecting voice activation in the microphone signal. In one example, the voice activation signal is input to a smoothing filter to limit the number of false positives of voice activation. In one example, the signal is processed frame by frame, and voice activation is indicated as a value for each frame, such as a binary value or a multi-level value. In one example, voice activation is detected only when a predetermined number of frames are determined for voice activation. In some examples, the predefined number of frames is at least 4 or 5 consecutive frames. Each frame can have a duration of approximately 30 milliseconds to 40 milliseconds, such as 33 milliseconds. Consecutive frames can have a time overlap of 40%-60% (e.g., 50%). This means that talk activation can be reliably detected in approximately 100 milliseconds or less or longer.
[0039] Typically, wearable devices can be configured as:
[0040] - headphones that can be worn on the wearer's head, for example by a headband, or around the wearer's neck, for example by a neckband;
[0041] - a pair of in-ear headphones that fit over the wearer's ears;
[0042] - A headset or a pair of earphones including one or more microphones and a transceiver to enable a headset mode of the headset or the pair of earphones.
[0043] Typically, headphones include earmuffs to sit above or on the wearer's ears, while in-ear headphones include earplugs or earplugs to be inserted into the wearer's ears. Herein, the earmuffs, earplugs or earplugs are referred to as earpieces. Earpieces are typically configured to establish a space between the eardrum and the speaker. A microphone may be arranged in the earpiece as an internal microphone to capture sound waves inside the space between the eardrum and the speaker, or may be arranged in the earpiece as an external microphone to capture sound waves impinging on the earpiece from the surrounding environment.
[0044] In some aspects, the microphone signal includes a first signal from an internal microphone. In some embodiments, the microphone signal includes a second signal from an external microphone. In some embodiments, the microphone signal includes a first signal and a second signal. The microphone signal may include one or both of the first signal and the second signal from a left side and a right side.
[0045] In some aspects, the processor is integrated into the body portion of the wearable device. The body portion may include one or more of the following: earpieces, headband, neckband, and other body portions of the wearable device. The processor may be configured as one or more components, such as having a first component in the left body portion of the wearable device and a second component in the right body portion.
[0046] In some aspects, the masking signal is received via a wireless or wired connection to an electronic device (eg, a smartphone or personal computer). The masking signal may be provided by an application (eg, an application including an audio player) running on the electronic device.
[0047] In some aspects, the microphone is a non-directional microphone, such as an omnidirectional microphone having a cardioid, supercardioid, or figure-8 characteristic.
[0048] In some embodiments, the processor is configured with one or both of the following:
[0049] - an audio player for generating a masking signal by playing an audio track; and
[0050] - an audio synthesizer for generating a masking signal using one or more signal generators.
[0051] Thus, a processor integrated in a wearable device may be configured with a player to generate a masking signal by playing an audio track. The audio track may be stored in a memory of the processor. The advantage is that the wearable device may be fully functional to transmit the masking signal without requiring a wired or wireless connection to the electronic device. This in turn may reduce power consumption, which is an advantage in relation to, for example, battery-powered electronic devices.
[0052] In some aspects, the audio track is uploaded from the electronic device to a memory of the wearable device as described above. In some aspects, the masking signal can be generated by the processor based on an audio stream or audio track received at the processor via a wireless transceiver at the wearable device. The audio stream or audio track can be transmitted by a media player at the electronic device such as a smart phone, tablet computer, personal computer, or server computer. The volume of the masking signal is controlled as described above.
[0053] The audio track may include, for example, audio samples according to a predefined codec. In some aspects, the audio track contains music, natural sounds, or a combination of artificial sounds that resemble one or more of music and natural sounds. The audio track may be selected, for example, from a predetermined set of audio tracks suitable for masking, via an application running on the electronic device. This allows the wearer to have more choices in terms of masking and to select or deselect certain tracks.
[0054] In some aspects, the player plays a track or a sequence of multiple tracks in an infinite loop.
[0055] In some aspects, a player is enabled to continuously play back a track or a sequence of multiple tracks at a time when a first criterion is met. The first criterion may be that the wearable device is in a first mode. In the first mode, the wearable device may be configured to function as a headset or an in-ear headset. The first criterion may additionally or alternatively include: the voice activation signal indicates voice activation. Thus, in accordance with the first criterion including the voice activation signal indicating voice activation, the player may resume playback in response to the voice activation signal transitioning from indicating that voice activation was not detected to indicating voice activation.
[0056] In some aspects, the synthesizer generates the masking by one or more noise generators generating colorful noise and by one or more modulators modifying the envelope of the signal from the noise generator. In some aspects, the synthesizer generates the masking signal according to stored instructions (e.g., MIDI instructions). The advantage is that the variation of the masking signal can be obtained by changing one or more parameters instead of the sampling sequence, which can reduce memory consumption while still providing flexibility.
[0057] In some embodiments, the processor is configured to include a machine learning component to generate a voice activation signal (y); wherein the machine learning component is configured to indicate a time period that the microphone signal includes:
[0058] - a signal component representing speech activation, or
[0059] - a signal component representing speech activity and a signal component representing noise (which is different from speech activity).
[0060] Thus, the machine learning component can be configured to achieve effective detection of voice activation and effective distinction between voice activation and voice inactivity.
[0061] The voice activation signal may be in the form of a time domain signal or a frequency time domain signal, for example represented by values arranged in a frame. The time domain signal may be a two-level or multi-level signal.
[0062] The machine learning component consists of a set of values encoded in one or both of hardware and software to indicate a time period. A set of values is obtained through a training process using training data. The training data may include input data recorded in a physical environment or synthesized based on, for example, a mixture of non-speech sounds and speech sounds. The training data may include output data indicating whether there is speech activation in the input data. The output data may be generated by an audio professional listening to an example of a microphone signal. Alternatively, in the case where the input data is synthesized, the output data may be generated by an audio professional or obtained from metadata or parameters used to synthesize the input data. The training data may be constructed or collected to include training data that at least primarily represents sound, such as sound from a selected sound source, from a predetermined acoustic environment (e.g., an office environment).
[0063] Examples of noises other than voice activation may be sounds from keys being pressed on a keyboard, sounds from an air conditioning system, sounds from a vehicle, etc. Examples of voice activation may be sounds from one or more people talking or shouting.
[0064] In some aspects, the machine learning component is characterized by indicating a likelihood that the microphone contains voice activation within a period of time.
[0065] In some aspects, the machine learning component is characterized to indicate the likelihood that the microphone signal contains voice activation and a signal component representing noise (which is different from voice activation for a period of time). For example, the signal component representing noise (different from voice activation) may come from keyboard keys.
[0066] The likelihood may be expressed in discrete form, such as binary form.
[0067] The machine learning component represents the correlation between:
[0068] - a voice activation signal with or without a noisy signal and a value representing the presence of voice activation; and
[0069] - a speech inactivity signal with and without a noise-based signal and a value representing the absence of speech activity;
[0070] This correlation is well known in the art.The microphone signal may include a voice active signal and a voice inactive signal.
[0071] In some aspects, the microphone signal is a frequency-time representation of an audio waveform in the time domain. In some aspects, the microphone signal is a frequency-time representation of an audio waveform in the time domain.
[0072] In some aspects, the machine learning component is a recursive neural network that receives samples of the microphone signal within a predefined window of the samples and outputs a voice activation signal. In some aspects, the machine learning component is a neural network, such as a deep neural network.
[0073] In some implementations, the machine learning component detects voice activation based on processing of the time domain waveform of the microphone signal.
[0074] The machine learning component can more effectively detect voice activation based on processing the time domain waveform of the microphone signal. This is particularly useful when other uses in the processor do not require frequency domain processing of the microphone signal.
[0075] In some aspects, the recurrent neural network has multiple input nodes that receive a sample sequence of a microphone signal, and at least one output node outputs a voice activation signal. The input node can receive the latest sample of the microphone signal. For example, the input node can receive the latest sample of the microphone signal corresponding to a window of duration of about 10 milliseconds (to 100 milliseconds, such as 30 milliseconds). The duration of the window can be shorter or longer.
[0076] As described above, in some aspects, the machine learning component is a neural network, such as a deep neural network. In some aspects, the machine learning component is a recurrent neural network, and voice activation is detected based on processing the time domain waveform of the microphone signal. Based on processing the time domain waveform of the microphone signal, the recurrent neural network may be more effective in detecting voice activation.
[0077] In some embodiments, the processor is configured to:
[0078] While receiving the microphone signal:
[0079] generating a frame comprising a frequency-time representation of a waveform of a microphone signal; wherein the frame comprises values arranged in frequency bins;
[0080] Included is a machine learning component configured to detect voice activation based on processing frames of a frequency-time representation of a waveform comprising a microphone signal.
[0081] When voice activation is present simultaneously with other noise activation signals, the machine learning component can more effectively detect voice activation based on processing frames of a frequency-time representation of a waveform containing the microphone signal.
[0082] In some aspects, the neural network is a recurrent neural network having a plurality of input nodes and at least one output node; wherein the processor is configured to:
[0083] 1) Inputting a sequence of all or part of the values in the selected frequency region into an input node of a recurrent neural network;
[0084] 2) outputting a corresponding voice activation signal for the selected frequency region at at least one output node; and
[0085] 3) Execute the above 1) and 2) simultaneously and / or sequentially for all or selected frequency regions of the frame.
[0086] In some embodiments, the neural network is a convolutional neural network having a plurality of input nodes and a plurality of output nodes. The plurality of input nodes may receive values of a frame according to a frequency-time representation and output values of the frame. In some aspects, the plurality of input nodes may receive values of a frame according to a time domain representation and output values of the frame.
[0087] Frames can be generated from overlapping sequences of samples of microphone signals. Frames can be generated from samples of approximately 30 milliseconds (e.g., including 512 samples). Frames can overlap each other by approximately 50%. Frames can include 257 frequency bins. Frames can be generated from longer or shorter sequences of samples. Likewise, the sampling rate can be faster or slower. The overlap can be greater or less.
[0088] The frequency-time representation may be in accordance with the MEL scale as described in: Stevens, Stanley Smith; Volkmann; John & Newman, Edwin B. (1937). "A scale for the measurement of the psychological magnitude pitch". Journal of the Acoustical Society of America. 8(3): 185–190. Alternatively, the frequency-time representation may be in accordance with an approximation thereof or in accordance with another scale having a logarithmic or approximately logarithmic relationship with the frequency scale.
[0089] The processor can be configured to generate frames comprising a frequency-time representation of a waveform of a microphone signal by one or more of: short-time Fourier transform, wavelet transform, bilinear time-frequency distribution function (Wigner distribution function), modified Wigner distribution function, Gabor-Wigner distribution function, Hilbert-Huang transform or other transforms.
[0090] In some embodiments, the machine learning component is configured to generate a voice activation signal based on a frequency-time representation comprising values arranged in frequency regions in a frame; wherein the processor controls the masking signal based on the time and frequency distribution of an envelope of the masking signal that substantially matches the voice activation signal (which is based on the frequency-time representation) or the envelope of the voice activation signal.
[0091] Thereby, the masking signal matches the speech activation, for example with respect to energy or power. This enables more accurate masking of the speech activation, which in turn can reduce the listening stress perceived by the wearer of the wearable device. The masking signal is different from the speech signal detected in the microphone signal. The masking signal is generated to mask the speech signal rather than to cancel the speech signal.
[0092] In some aspects, the processor is configured to generate a masking signal by mixing a plurality of intermediate masking signals; wherein the processor controls one or both of the mixing and content of the intermediate masking signals to have a time and frequency distribution that matches the voice activation signal (which is represented according to a frequency-time representation). The processor may also synthesize the masking signal as described above so that the time and frequency distribution matches the voice activation signal.
[0093] Thus, the masking signal may be composed to match the energy level of the microphone signal in frequency bands determined to contain speech activity.In frequency bands determined to contain speech inactivity, the masking signal is composed to not match the energy level of the microphone signal.
[0094] In some embodiments, the processor is configured to:
[0095] In response to detecting an increase in the frequency or density of voice activity, the volume of the masking signal is gradually increased over time.
[0096] Thus, a good compromise can be achieved between early masking at the onset of speech activation and the reduction of auditory artifacts due to the masking signal.
[0097] In some aspects, the processor is configured to gradually reduce the volume of the masking signal over time in response to detecting a decrease in the frequency or density of voice activation. Thus, the masking signal decays rather than being abruptly disconnected or turned off. In particular, the risk of audible artifacts that may be offensive to the device wearer is reduced.
[0098] In some embodiments, the processor is configured with:
[0099] A mixer for generating a masking signal from one or more intermediate masking signals selected from a plurality of intermediate masking signals; wherein the one or more selected intermediate masking signals are selected based on criteria based on one or both of a microphone signal and a voice activation signal.
[0100] Thus, the masking signal can be configured from a plurality of possible combinations. In some aspects, the mixer is configured with mixer settings. The mixer settings can include gain settings for each intermediate masking signal.
[0101] In some embodiments, the processor is configured with:
[0102] a gain stage configured with a trigger for attacking amplitude modulation of the intermediate masking signal and a trigger for attenuating amplitude modulation of the intermediate masking signal;
[0103] wherein, in response to detecting a transition from speech inactivity to speech activity, the gain stage is triggered to perform enhanced amplitude modulation of the intermediate masking channel, and in response to detecting a transition from speech activity to speech inactivity, the gain stage is triggered to perform attenuated amplitude modulation of the intermediate masking channel.
[0104] Thereby, artifacts in the masking signal due to processing of the masking signal may be kept at an inaudible level or reduced.In some aspects, multiple intermediate masking signals are generated simultaneously or sequentially through multiple gain stages.The intermediate masking signals may be mixed as described above.
[0105] In some embodiments, the processor is configured with:
[0106] an active noise reduction unit for processing the microphone signal and providing an active noise reduction signal to the speaker; and
[0107] A mixer is used to mix the active noise reduction signal and the masking signal into a signal for the loudspeaker.
[0108] In particular, active noise cancellation (ANC) is effective in cancelling out tonal noise, such as from machinery. However, this makes voice activation less intelligible and more disruptive to the wearer of the wearable device. However, combined with masking applied when voice activation is detected, the improvement in the sound environment perceived by the wearer exceeds active noise cancellation and goes beyond masking.
[0109] In some aspects, active noise reduction is achieved through a feedforward configuration, a feedback configuration, or through a hybrid configuration. As described above, in the feedforward configuration, the wearable device is configured with an external microphone. The external microphone forms a reference noise signal for the ANC algorithm. In the feedback configuration, as described above, an internal microphone is placed to form a reference noise signal for the ANC algorithm. The hybrid configuration combines the feedforward and feedback configurations and requires at least two microphones arranged in the feedforward and feedback configurations, respectively.
[0110] The microphone used to generate the microphone signal used to generate the masking signal may be an internal microphone or an external microphone.
[0111] In some embodiments, the processor is configured to selectively operate in a first mode or a second mode;
[0112] wherein, in the first mode, the processor controls the volume of the masking signal provided to the speaker; and
[0113] Among them, in the second mode, the processor:
[0114] - regardless of whether the voice activation signal indicates voice activation, ceasing to provide the masking signal to the speaker at the first volume.
[0115] In this way, in the second mode, for example, when the wearer speaks to a voice recorder coupled to receive a microphone signal, speaks to a digital assistant coupled to receive a microphone signal, speaks to a remote party coupled to receive a microphone signal, or speaks to a person near the wearer while wearing the wearable device, the masking signal does not interfere with the wearer.
[0116] In some aspects, in a first mode, the wearable device functions as a headset or earphones. The first mode can be a focus mode in which active noise reduction and / or active reduction of conversation intelligibility is applied by masking signals. In a second mode, the wearable device is used as a headset. When enabled to function as a headset, the wearable device can participate in a call with a remote party of a call.
[0117] The second mode may be selected by activating an input mechanism such as a button on the wearable device.The first mode may be selected by activating or reactivating an input mechanism such as a button on the wearable device.
[0118] In some aspects, the processor stops providing the masking signal to the speaker in the second mode or provides the masking signal to the speaker at a low volume without disturbing the wearer. In some aspects, in the second mode, the processor stops enabling or disabling providing the masking signal to the speaker.
[0119] Thus, a wearable device may be configured with a hear-through mode that is selectively enabled by a user of the wearable device.
[0120] In some embodiments, the electroacoustic input transducer is a first microphone that outputs a first microphone signal; and wherein the wearable device comprises:
[0121] - a second microphone, which outputs a second microphone signal; and
[0122] A beamformer coupled to receive the first microphone signal or the third microphone signal and the second microphone signal from the third microphone and to generate a beamformed signal.
[0123] In some aspects, in the second mode defined above, the beamforming signal is provided to a transmitter, which is engaged to transmit a signal based on the beamforming signal to a remote receiver.
[0124] The beamformer may be an adaptive beamformer or a fixed beamformer. The beamformer may be a broadside beamformer or an endfire beamformer.
[0125] A signal processing method on a wearable electronic device is also provided, the wearable electronic device comprising: an electroacoustic input transducer arranged to pick up an acoustic signal and convert the acoustic signal into a microphone signal; a speaker; and a processor, which performs the following operations:
[0126] controlling the volume of the masking signal; and
[0127] providing a masking signal to a loudspeaker;
[0128] Based on processing at least the microphone signal, detecting voice activation and generating a voice activation signal concurrently with the microphone signal, the voice activation signal sequentially indicating one or more of: voice activation and voice inactivity; and
[0129] In response to the voice activation signal, a volume of the masking signal is controlled according to providing the masking signal to the speaker at a first volume when the voice activation signal indicates voice activation and providing the masking signal to the speaker at a second volume when the voice activation signal indicates voice inactivity.
[0130] Various aspects of the method are defined in the general description and in the dependent claims in conjunction with a wearable device.
[0131] A signal processing module for a headphone or earphone is also provided, which is configured to perform the method.
[0132] The signal processing module may be a signal processor, for example in the form of an integrated circuit or a plurality of integrated circuits arranged on one or more circuit boards or a part thereof.
[0133] Also provided is a computer readable medium comprising instructions for performing the method when executed by a processor at a wearable electronic device, the wearable electronic device comprising an electroacoustic input transducer arranged to pick up an acoustic signal and convert the acoustic signal into a microphone signal; and a speaker.
[0134] The computer readable medium may be the memory of the signal processing module or a part thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0135] A more detailed description is given below with reference to the accompanying drawings, in which:
[0136] Figure 1 A wearable electronic device embodied as a headset and a pair of in-ear headphones and a block diagram of the wearable device are shown;
[0137] Figure 2 A module for generating a masking signal is shown, the module comprising an audio player;
[0138] Figure 3 A module for generating a masking signal is shown, the module comprising an audio synthesizer;
[0139] Figure 4 shows a spectrogram of a microphone signal and a corresponding spectrogram of a voice activation signal;
[0140] Figure 5 A gain stage is shown, which is configured with a flip-flop for amplitude modulation of the masking signal; and
[0141] Figure 6 A block diagram of a wearable device with a headphone mode and an earphone mode is shown. DETAILED DESCRIPTION
[0142] Figure 1 A wearable electronic device embodied as a headset or a pair of in-ear headphones and a block diagram of the wearable device are shown.
[0143] The headset 101 includes a headband 104 carrying a left earpiece 102 and a right earpiece 103, which may also be referred to as ear cups. A pair of in-ear headphones 116 includes a left earpiece 115 and a right earpiece 117.
[0144] The earpiece includes at least one speaker 105, such as a speaker in each earpiece. The headset 101 also includes at least one microphone 106 in the earpiece. As described herein, hereinafter, a headset or a pair of in-ear headphones may include a processor configured in a selectable headphone mode in which masking is disabled or significantly reduced.
[0145] The block diagram of the wearable device shows an electroacoustic input transducer in the form of a microphone 106 (arranged to pick up an acoustic signal and convert the acoustic signal into a microphone signal x), a speaker 105 and a processor 107. The microphone signal may be a digital signal, or may be converted into a digital signal by a processor. The speaker 105 and the microphone 105 are generally referred to as an electroacoustic transducer element 114. The electroacoustic transducer element 114 of the wearable electronic device may include at least one speaker in the left-hand earpiece and at least one speaker in the right-hand earpiece. The electroacoustic transducer element 114 may also include one or more microphones arranged in one or both of the left-hand earpiece and the right-hand earpiece. The arrangement of the microphone in the right-hand earpiece may be different from that in the left-hand earpiece.
[0146] The processor 107 includes a voice activation detector VAD 108 that outputs a voice activation signal y (which may be a time domain voice activation signal or a frequency-time domain voice activation signal). The voice activation signal y is received by a gain stage G 110, which sets a gain factor in response to the voice activation signal. The gain stage may have two or more gain factors (e.g., multiple gain factors) that are selectively set in response to the voice activation signal. The gain stage G 110 may also be controlled in response to a microphone signal (e.g., via a filter or a circuit that implements adaptive gain control of a masking signal according to a feedforward or feedback configuration). The masking signal m may be generated by a masking signal generator 109. The masking signal generator 109 may also be controlled by the voice activation signal y. The masking signal m may be provided to the speaker 105 via a mixer 113. The mixer 113 mixes the masking signal m with a noise reduction signal q. The noise reduction signal is provided by a noise reduction unit ANC 112. The noise reduction unit ANC 112 may receive a microphone signal x from the microphone 106 and / or another microphone signal from another microphone arranged at a different location in the headphone or in-ear headphones than the microphone 106. The masking signal generator 109, the voice activation detector 108 and the gain stage 110 may be constituted by a signal processing module 111.
[0147] Thus, the processor 107 is configured to detect voice activation in the microphone signal and generate a voice activation signal y, which in turn indicates at least one or more of voice activation and voice inactivation. In addition, the processor 107 is configured to control the volume of the masking signal m in response to the voice activation signal y according to providing the masking signal m to the speaker 105 at a first volume when the voice activation signal y indicates voice activation, and providing the masking signal m to the speaker 105 at a second volume when the voice activation signal y indicates voice inactivation. The first volume can be controlled in response to the energy level or envelope of the microphone signal or the energy level or envelope of the voice activation signal. The second volume can be enabled by not providing the masking signal to the speaker or by controlling the volume to be about 10 dB or less below the microphone signal.
[0148] Also shown is a graph 118 showing that the gain factor of gain stage G 110 is relatively high when the voice activity signal indicates voice activity (va) and relatively low when the voice activity signal indicates voice inactivity (vi-a). The gain factor may be controlled in two or more steps.
[0149] Figure 2A module for generating a masking signal is shown, which includes an audio player. Module 111 includes a voice activation detector 108 and an audio player 201 and a gain stage G 110. The audio player 201 is configured to play an embedded audio track 202 or an external audio track 203. The audio track 202 or 203 may include encoded audio samples, and the player may be configured with a decoder for generating an audio signal from the encoded audio samples. The advantage of the embedded audio track 202 is that the wearable device can be configured with the audio track once or in response to a predetermined event. The embedded audio track can then be played without establishing a wired or wireless connection with a remote server or other electronic device; this in turn can save battery power for battery-powered wearable devices. The advantage of the external audio track 203 is that the content of the audio track can be changed according to preferences or predefined events. The voice activation detector 108 can send a signal y' to the player 201. The signal y' may convey a play command after detecting voice activation and convey a "stop" or "pause" command when voice inactivity is detected.
[0150] Figure 3 A module for generating a masking signal is shown, which includes an audio synthesizer. Module 111 includes a voice activation detector 108, an audio synthesizer 301 and a gain stage G 110. The synthesizer 301 can generate a masking signal according to a parameter 302. The parameter 302 can be defined by hardware or software, and in some embodiments can be selected according to the voice activation signal y. The synthesizer 301 includes one or more tone generators 305, 306 coupled to corresponding modulators 303, 304 (which can modulate the dynamics of the signal from the tone generator 305, 306). The modulators 303, 304 can operate according to the parameters 302. The modulators 303, 304 output intermediate masking signals m" and m"', which are input to a mixer 307, which mixes the intermediate masking signals to provide the masking signal m' to the gain stage 110. Modulation of the dynamics of the signal from the tone generators 305, 306 may change the envelope of the signal from the tone generator(s).
[0151] Although volume control is described with respect to gain stage G 110, it should be noted that volume control may be implemented in other ways, such as by controlling the modulation or generation of the content of the masking signal itself.
[0152] Figure 4 A spectrogram of a microphone signal and a corresponding spectrogram of a voice activation signal are shown. In general, a spectrogram is a spectrum visual representation of the frequency of a signal varying over time. The spectrogram is displayed along a time axis (horizontally) and a frequency axis (vertically). The spectrogram shown as an illustrative example spans a frequency range of approximately 0 Hz to 8000 Hz and a time period of approximately 0 seconds to 10 seconds.
[0153] The spectrogram 401 (left hand panel) of the microphone signal includes a first region 403 where the signal energy is distributed over a wide frequency range and occurs at about 2-3 seconds. The signal energy ranges up to 0 dB and comes mainly from the keys on the keyboard.
[0154] The second region 404 contains signal energy that is distributed over a wider frequency range within a range below about -20 dB and occurs at about 4-6 seconds. This signal energy comes primarily from indistinguishable noise sources, sometimes referred to as background noise.
[0155] The third region represents the presence of a conversation in the microphone signal and includes a first portion 407 representing the most dominant part of the conversation at lower frequencies and a second portion 405 representing a less dominant part of the conversation at higher frequencies over a wider frequency range. The conversation occurs at approximately 7-8 seconds.
[0156] The output of a voice activation detector (e.g., voice activation detector 108) is shown in the spectrogram 402 (right hand panel). As can be seen, the output of the voice activation detector is also located at about 7-8 seconds. The output level of the voice activation detector corresponds to the energy level of the talk signal, with a larger dominant portion 408 at lower frequencies and a smaller dominant portion 406 at higher frequencies over a wider frequency range.
[0157] The output of the voice activity detector is thus shown as a spectrogram according to the corresponding frame representation. The output of the voice activity detector is used to control the volume of the masking signal and optionally generate the content of the masking signal according to the desired spectral distribution. The output of the voice activity detector can be reduced to a one-dimensional binary or multi-level signal time domain signal without spectral decomposition.
[0158] Figure 5 A gain stage 501 is shown which is configured with a trigger for amplitude modulation of the masking signal.This embodiment is an example of how to adapt the masking signal based on the voice activation signal y to obtain a desired fade-in and / or fade-out of the masking signal m.
[0159] The first trigger unit 505 detects the start of speech activity, for example by a threshold, and activates the fade-in modulation characteristic 503. The modulator 502 applies the fade-in modulation characteristic 503 to modulate the intermediate masking signal m'' to generate another intermediate masking signal m', which is provided to the gain stage G110.
[0160] The second trigger unit 506 detects the end or decrease of the voice activation period, for example by a threshold, and activates the fade-out modulation characteristic 504. The modulator 502 applies the fade-out modulation characteristic 504 to modulate the intermediate masking signal m'' to generate another intermediate masking signal m', which is provided to the gain stage G110.
[0161] Thereby, artifacts in the masking signal may be reduced.
[0162] Figure 6 A block diagram of a wearable device having a headphone mode and a headset mode is shown. In some aspects, the block diagram corresponds to the above block diagram, but also includes elements included in the headset module 601 related to enabling the headset mode. In addition, a selector 605 is provided for selectively enabling the headset mode or the headphone mode. The selector 605 can provide a masking signal m or a headphone signal f to the speaker 105. The selector can be engaged with other elements of the processor or separated therefrom. The headset block 601 may include a beamformer 602 that receives a microphone signal x from a microphone 106 and another microphone signal x' from another microphone 106'. The beamformer can be a broadside beamformer or an end-fire beamformer or an adaptive beamformer. The beamformed signal is output from the beamformer and is provided to a transceiver 604 that provides wired or wireless communication with an electronic communication device 606 such as a mobile phone or a computer.
[0163] Generally, it should be noted that, as known in the art, headphones or earphones may include elements for playing music. In this regard, playing music for the purpose of listening to music can be achieved through mode selection, which disables the masking of the above-mentioned voice activation control.
[0164] Generally, it should be understood that those skilled in the art can perform experiments, investigations and measurements to obtain an appropriate volume level for the masking signal. In addition, experiments, investigations and measurements may be required to avoid (non-linear) signal processing associated with the masking signal introducing audible or disturbing artifacts.
Claims
1. A wearable electronic device (101), comprising: an electroacoustic input transducer (106) arranged to pick up an acoustic signal and convert the acoustic signal into a microphone signal (x); Speaker (105); and The processor (107) is configured to: controlling a volume of a masking signal (m), wherein the masking signal is configured to mask ambient voice activation; and providing the masking signal (m) to the speaker (105); Characterized in that the processor is further configured to: Based on processing at least the microphone signal (x), detecting voice activation and generating a voice activation signal (y) concurrent with the microphone signal, the voice activation signal sequentially indicating one or more of: voice activation and voice inactivity; and In response to the voice activation signal (y), controlling the volume of the masking signal (m) according to providing the masking signal (m) to the speaker (105) at a first volume when the voice activation signal (y) indicates voice activation and providing the masking signal (m) to the speaker (105) at a second volume when the voice activation signal (y) indicates voice inactivity, wherein the voice activation signal is processed frame by frame and the voice activation is indicated as a value for each frame, the voice activation being determined to be detected only when a predetermined number of frames have been determined for the voice activation, The processor is configured with one or both of the following: - an audio player (201) that generates the masking signal by playing an audio track; and - an audio synthesizer (111) generating the masking signal using one or more signal generators.
2. The wearable electronic device according to claim 1, wherein: The processor is configured to include a machine learning component to generate the voice activation signal (y); wherein the machine learning component is configured to indicate a time period during which the microphone signal (x) includes: - a signal component representing speech activation, or - a signal component representing speech activity and a signal component representing noise other than speech activity.
3. The wearable electronic device according to claim 1, wherein: The machine learning component is configured to detect the voice activation based on processing of a time domain waveform of the microphone signal (x).
4. The wearable electronic device according to claim 1, wherein: The processor is configured to: While receiving the microphone signal: generating a frame comprising a frequency-time representation (X) of a waveform of the microphone signal (x); wherein the frame comprises values arranged in frequency bins; A machine learning component is included that is configured to detect the voice activation based on processing the frame comprising a frequency-time representation of a waveform of the microphone signal (x).
5. The wearable electronic device according to claim 3 or 4, in, The machine learning component is configured to generate the voice activation signal (y) from a frequency-time representation comprising values arranged in frequency bins in a frame; Therein, the processor (107) controls the masking signal (m) according to the frequency-time representation according to a time and frequency distribution of the envelope of the masking signal that substantially matches the voice activation signal or an envelope of the voice activation signal.
6. The wearable electronic device according to any one of claims 1 to 3, wherein: The processor is configured to: In response to detecting an increase in the frequency or density of voice activity, gradually increasing the volume of the masking signal (m) over time.
7. The wearable electronic device according to any one of claims 1 to 3, wherein: The processor (107) is configured with: A mixer generates the masking signal based on selected one or more intermediate masking signals from a plurality of intermediate masking signals; wherein the selection of the selected one or more intermediate masking signals is performed based on a criterion based on the microphone signal and / or the voice activation signal.
8. The wearable electronic device according to any one of claims 1 to 3, wherein: The processor is configured with: a gain stage configured with a trigger for enhancing amplitude modulation of the intermediate masking signal and a trigger for attenuating amplitude modulation of the intermediate masking signal; wherein, in response to detecting a transition from speech inactivity to speech activity, the gain stage is triggered to perform enhanced amplitude modulation of the intermediate masking channel, and in response to detecting a transition from speech activity to speech inactivity, the gain stage is triggered to perform attenuated amplitude modulation of the intermediate masking channel.
9. The wearable electronic device according to any one of claims 1 to 3, wherein: The processor is configured with: an active noise reduction unit (112) for processing the microphone signal (x) and providing an active noise reduction signal (q) to the speaker; and A mixer (113) mixes the active noise reduction signal (q) and the masking signal (m) into a signal for the speaker (105).
10. The wearable electronic device according to any one of claims 1 to 3, wherein: The processor (107) is configured to selectively operate in a first mode or a second mode; wherein, in the first mode, the processor (107) controls the volume of the masking signal (m) provided to the speaker (105); and Wherein, in the second mode, the processor (107): - ceasing to provide the masking signal (m) to the loudspeaker (105) at the first volume independently of the voice activation signal (y) indicating voice activation.
11. The wearable electronic device according to any one of claims 1 to 3, wherein: The electroacoustic input transducer is a first microphone (106) outputting a first microphone signal (x); and wherein the wearable electronic device comprises: - a second microphone (106') outputting a second microphone signal (x'); and A beamformer coupled to receive a third microphone signal from a third microphone or the first microphone signal (x), and the second microphone signal (x'), and to generate a beamformed signal.
12. A signal processing method at a wearable electronic device (101), the wearable electronic device comprising: An electroacoustic input transducer (106) arranged to pick up an acoustic signal and convert said acoustic signal into a microphone signal (x); a loudspeaker (105); and a processor (107) for performing the following: controlling a volume (m) of a masking signal, wherein the masking signal is configured to mask ambient voice activation; and providing the masking signal (m) to the speaker (105); Based on processing at least the microphone signal (x), detecting voice activation and generating a voice activation signal (y) concurrent with the microphone signal, the voice activation signal sequentially indicating one or more of: voice activation and voice inactivity; and In response to the voice activation signal (y), controlling the volume of the masking signal (m) according to providing the masking signal (m) to the speaker (105) at a first volume when the voice activation signal (y) indicates voice activation and providing the masking signal (m) to the speaker (105) at a second volume when the voice activation signal (y) indicates voice inactivity, wherein the voice activation signal is processed frame by frame and the voice activation is indicated as a value for each frame, the voice activation being determined to be detected only when a predetermined number of frames have been determined for the voice activation, The processor is configured with one or both of the following: - an audio player (201) that generates the masking signal by playing an audio track; and - an audio synthesizer (111) generating the masking signal using one or more signal generators.
13. A computer readable medium comprising instructions for executing the signal processing method according to claim 12 when executed by a processor (107) at a wearable electronic device (101), the wearable electronic device (101) comprising: an electroacoustic input transducer (106) arranged to pick up an acoustic signal and convert the acoustic signal into a microphone signal (x); Speaker (105).
Citation Information
Patent Citations
Noise Masking in Headsets
US20150348530A1
Adapted audio masking
US8964997B2
Adapted Audio Masking
US20110235813A1
Dynamically adjustable sidetone generation
US20190306608A1
System and method to detect close voice sources and automatically enhance situation awareness
US9270244B2