Low latency automatic mixer with integrated voice and noise activity detection
By combining a voice activity detector and a channel gate control module, and dynamically adjusting channel gate control decisions, the problems of erroneous noise rejection and reduced signal-to-noise ratio in the conference environment are solved, achieving low-latency, high-signal-to-noise ratio audio output and improving user satisfaction.
Patent Information
- Application Number
- CN202080048155.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-31
- Filing Date
- 2020-05-29
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2040-05-29
AI Technical Summary
Existing technologies struggle to effectively filter out non-speech or non-human noise in conference and presentation environments, leading to reduced signal-to-noise ratio, increased latency, and decreased user satisfaction. Furthermore, existing systems cannot effectively address the spatial relationship between speech and noise caused by imperfect acoustic polarity patterns of microphones and beamlobes.
By combining a voice activity detector and a channel gate control module, the channel gate control decision is dynamically adjusted. The noise activity detector identifies erroneous noise and shuts down the relevant channels. Combined with front-end noise leakage minimization technology, the contribution of erroneous noise is reduced, while maintaining low-latency audio output.
It effectively eliminates erroneous noise under low latency conditions, improves the signal-to-noise ratio and user satisfaction, reduces the impact of front-end noise leakage on audio quality, and enhances the accuracy of microphone selection and beam lobe selection.
Smart Images

Figure CN114051637B_ABST
Abstract
Description
[0001] Cross-citation of related applications
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 855,491, filed May 31, 2019, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application generally relates to systems and methods for providing low-latency speech and noise activity detection integrated with an automatic audio mixer. Specifically, this application relates to systems and methods for providing speech and noise activity detection using an automatic audio mixer that excludes erroneous non-speech or non-human noise while maximizing the signal-to-noise ratio and minimizing audio latency. Background Technology
[0004] Meeting and presentation environments (such as boardrooms, conference settings, and the like) may involve using multiple microphones or microphone array lobes to capture sound from various audio sources. For example, the audio sources may include human speakers. The captured sound can be amplified by loudspeakers (for sound reinforcement) and propagated to a local audience in the environment and / or others far from the environment (e.g., via television broadcasting and / or internet broadcasting). Each of the microphones or array lobes can form a channel. The captured sound can be provided as a multi-channel audio input and as a single mixed audio channel.
[0005] Typically, captured sound may also contain erroneous non-speech or non-human noise from the environment, such as sudden, impactful, or recurring sounds like page turning, opening packages and containers, chewing, typing, etc. To minimize erroneous noise in the captured sound, speech activity detection (VAD) algorithms and / or automatic mixers can be applied to the channels of microphones or array lobes. Automatic mixers automatically reduce the strength of the audio input signal of a particular microphone to mitigate the contribution of background, static, or stationary noise when it is not capturing human speech or voice. VAD is a technique used in speech processing where the presence or absence of human speech or voice is detected. Additionally, noise reduction techniques can reduce certain background, static, or stationary noises, such as fan and HVAC system noise. However, such noise reduction techniques are not ideal for reducing or eliminating erroneous noise.
[0006] While combinations of automatic mixing and VAD exist in current systems, such combinations typically fail to inherently exclude error noise (specifically, low audio delays that enable real-time communication or use with room sound reinforcement). This exclusion of error noise can impair the performance of typical automatic mixers, which often rely on relatively simple channel selection rules, such as first arrival time or highest amplitude at a given moment. Current systems integrating automatic mixing and VAD may not be optimal due to high latency and / or front-end clipping (FEC) in speech or audio. For example, additional audio delay could be added to the channel to align the VAD detection delay to speech occurrence, minimizing FEC of syllables or words in the speech or audio stream, but this could result in unacceptable latency in the audio stream. Alternatively, FEC could be accepted by deciding not to add audio delay to align the VAD detection delay to the audio stream, but this could result in incomplete speech or audio in the audio stream. These conditions can lead to reduced user satisfaction. Furthermore, many current systems with VADs can utilize only a single audio channel, where effective operation does not require consideration of the spatial relationship between speech / voice and noise occurring in a specific environment.
[0007] Furthermore, in automatic mixing applications (with individual microphone units or using manipulated audio lobes from a microphone array), due to imperfect acoustic polar patterns of the microphones and / or lobes, speech and error noise can occur in the same environment and can be contained in all microphones and / or lobes. This can lead to problems with VAD detection capabilities (on both individual and common channels), appropriate automatic mixer channel selection (which attempts to avoid error noise while still selecting channels containing speech), and suppressing error noise in lobes that are gated open due to the presence of speech / voice.
[0008] Therefore, the system and method have the potential to address these issues. More specifically, the system and method have the potential to provide speech and noise activity detection using an automatic audio mixer that can exclude erroneous non-speech or non-human noise while maximizing the signal-to-noise ratio, increasing intelligibility, minimizing audio latency, and increasing user satisfaction. By combining automatic mixing principles with more advanced speech activity detection techniques, microphone / lobe selection can be enhanced to maximize the speech-to-noise ratio. Summary of the Invention
[0009] The present invention aims to solve the problems mentioned above by providing systems and methods, which are particularly designed to: (1) utilize a modified speech activity detector, which is modified to function as a noise activity detector to sense the presence of speech or erroneous noise on a channel; (2) perform additional channel gate control based on metrics and decisions from the speech activity detector, which may influence and / or modify channel gate control performed by an automatic mixer; (3) reduce or eliminate the amount of front-end clipping of captured speech / voice; and (4) minimize the effect of front-end noise leakage from erroneous noise that may initially be included in a particular gate-open channel.
[0010] In one embodiment, a method includes: determining whether non-voice audio exists in an audio signal of a channel initially gated open by a mixer, wherein the mixer generates a mixed audio signal based on at least the audio signal of the initially gated open channel; and, upon determining that the non-voice audio exists in the audio signal of the initially gated open channel, modifying the mixer by gate-closing the initially gated open channel to cause the mixer to generate the mixed audio signal that does not have the audio signal of the initially gated open channel.
[0011] In another embodiment, a system includes an activity detector configured to determine whether non-voice audio is present in the audio signal of a channel initially gated open by a mixer, wherein the mixer is configured to generate a mixed audio signal based on at least the audio signal of the initially gated open channel. The system further includes a channel gate module communicating with the activity detector, and the channel gate module is configured to modify the mixer to gate the initially gated open channel and generate the mixed audio signal that does not contain the audio signal of the initially gated open channel when the activity detector determines that the non-voice audio is present in the audio signal of the initially gated open channel.
[0012] These and other embodiments, as well as various substitutions and aspects, will be understood and more fully appreciated from the following detailed description and accompanying drawings, which illustrate various ways in which the principles of the invention can be employed. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of a system including a mixer and a voice activity detector for channel gating, according to some embodiments.
[0014] Figure 2 This describes the use of some embodiments. Figure 1 The flowchart shows the system gate operation from the microphone channel.
[0015] Figure 3Based on some embodiments Figure 1 A diagram of an exemplary gated state machine used in the mixer of the system. Detailed Implementation
[0016] The following description illustrates, describes, and exemplifies one or more specific embodiments of the invention based on the principles of the invention. This description is not intended to limit the invention to the embodiments described herein, but rather to illustrate and teach the principles of the invention so that those skilled in the art can understand these principles and, using this understanding, apply them to practice not only the embodiments described herein, but also other embodiments conceivable based on these principles. The scope of the invention is intended to cover all such embodiments that, literally or according to the principle of equivalents, fall within the scope of the appended claims.
[0017] It should be noted that similar or substantially similar elements may be designated using the same element symbols in the description and accompanying drawings. However, these elements may sometimes be designated using different numerical symbols, for example, where such designation is advantageous for clearer description. Furthermore, the drawings illustrated herein are not necessarily drawn to scale, and in some instances, the scale may have been exaggerated to more clearly depict certain features. Such designation and graphical practices do not necessarily imply any underlying material purpose. As stated above, this specification is intended to be considered holistically and according to the principles of the invention as taught herein and understood by one of ordinary skill in the art.
[0018] The systems and methods described herein generate mixed audio signals from an automatic mixer that reduces and minimizes the contribution of erroneous non-speech or non-human noise sensed in the environment. The systems and methods may utilize the automatic mixer in conjunction with a speech activity detector (or an erroneous noise activity detector), each making independent channel gating decisions. The automatic mixer may gating specific channels open or closed based on channel selection rules, while the speech / erroneous noise activity detector may modify the automatic mixer's channel gating decisions depending on whether speech or erroneous noise is detected in channels gated open by the automatic mixer. Metrics from the speech / erroneous noise activity detector (e.g., confidence scores) may also influence channel gating decisions and / or the relative selection of each channel in the automatic mixer. To support low-latency audio output, some erroneous noise may leak into the audio mix before the speech / erroneous noise activity detector can modify the audio mixer. The systems and methods allow this behavior while minimizing the energy and subjective audio quality impact of this channel gating noise initiation. This allows minimizing the energy from erroneous noise leaking into the channels while maintaining low latency.
[0019] Figure 1This is a schematic diagram of a system 100 that can be used to reject erroneous noise, which includes a microphone 102, a mixer 104, and a voice activity detector 108. Figure 2 It is for use Figure 1 The flowchart of the process 200 for system 100 to exclude erroneous noise is provided. System 100 and process 200 can produce an output of a mixed audio signal with optimal signal-to-noise ratio and containing the desired speech, while minimizing the inclusion or contribution of erroneous noise.
[0020] For example, a conference room environment can utilize System 100 to facilitate communication with personnel at remote locations. The type of microphone 102 and its placement in a particular environment may depend on the location of the audio source, physical space requirements, aesthetics, room layout, and / or other considerations. For example, in some environments, the microphone may be placed on a table or lectern near the audio source. In other environments, the microphone may be mounted overhead to capture sound from across the room. System 100 can work with any type and any number of microphones 102. The various components included in System 100 may be implemented using software that can be executed by one or more servers or computers (e.g., computing devices with processors and memory, graphics processing units (GPUs)) and / or by hardware (e.g., discrete logic circuits, application-specific integrated circuits (ASICs), programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.).
[0021] Generally, a computer program product according to an embodiment includes a computer-usable storage medium (e.g., standard random access memory (RAM), optical disc, universal serial bus (USB) flash drive, or the like) having computer-readable program code embodied therein, wherein the computer-readable program code is adapted to be executed by a processor (e.g., working with an operating system) to implement the methods described below. In this regard, the program code can be implemented in any desired language and can be implemented as machine code, combinatorics, bytecode, interpretable source code, or the like (e.g., via C, C++, Java, Actionscript, Objective-C, JavaScript, CSS, XML, and / or others).
[0022] refer to Figure 1System 100 may include microphone 102, mixer 104, premixer 106, voice activity detector 108, and channel gate module 110. Each of the microphones 102 can detect ambient sound and convert the sound into an audio signal to form a channel. In embodiments, some or all of the audio signals from the microphones 102 may be processed by a beamformer (not shown) to generate one or more beamformed audio signals, as known in the art. Therefore, although the system and method described herein are based on audio signals from microphone 102, it is contemplated that the system and method may also utilize any type of sound source, such as beamformed audio signals generated by a beamformer.
[0023] Audio signals from each of the microphones 102 can be received by mixer 104, premixer 106, and voice activity detector 108, for example in Figure 2 Step 202 of process 200 is shown in the diagram. Mixer 104 can ultimately generate and output a mixed audio signal that conforms to a desired audio mix, such that audio signals from certain microphones are emphasized and deemphasized or suppressed from other microphones. Exemplary embodiments of audio mixers are disclosed in commonly assigned patents (U.S. Patent Nos. 4,658,425 and 5,297,210), the entire contents of each of which are incorporated herein by reference.
[0024] The mixed audio signal from mixer 104 may include contributions from one or more channels gated open using system 100, i.e., audio signals from microphone 102. Mixer 104 and channel gate module 110 may gate one or more channels to provide captured audio without suppression (or, in some embodiments, with minimal suppression) in response to determining that the captured audio contains human speech and / or according to certain channel selection rules. Mixer 104 and channel gate module 110 may also gate one or more channels to reduce the intensity of some captured audio in response to determining that the captured audio in a channel is background, static, or fixed noise. Channel gate determination by mixer 104 and channel gate module 110 may occur in step 204. Mixer 104 and channel gate module 110 may present a channel gate decision for each of the multiple channels corresponding to the multiple microphones 102 (or array lobes). Process 200 may continue to step 206.
[0025] In step 206, if the channel was determined to be gated off in step 204, then process 200 may continue to step 218 and mixer 104 may output a mixed audio signal that does not contain the gated-off channel. However, in step 206, if the channel was determined to be gated on in step 204, then process 200 may continue to step 208, where in some embodiments a non-voice deemphasis filter may be applied, which serves as a bandwidth-limiting filter (e.g., a low-pass filter, a band-pass filter, or linear predictive coding (LPC)) to subjectively minimize front-end noise leakage, as described in further detail below.
[0026] In step 210, the Voice Activity Detector (VAD) 108 may also receive audio signals from the microphone 102. The VAD 108 may execute an algorithm in step 210 to determine whether speech is present in a particular channel or conversely, whether noise is present in a particular channel. For example, if the VAD 108 detects speech in a particular channel (or no noise is detected), then the VAD 108 may consider the channel to contain speech or be "non-noise". Similarly, if the VAD 108 does not detect speech in a particular channel (or noise is detected), then the channel may be considered to contain noise or be "non-speech". In embodiments, the VAD 108 may be implemented by analyzing the spectral variance of the audio signal, using linear predictive coding (LPC), applying machine learning or deep learning techniques to detect speech, and / or using well-known techniques such as ITU G.729 VAD, the ETSI standard for VAD calculation included in the GSM specification, or long-range interval prediction.
[0027] By identifying whether a specific channel contains erroneous noise (i.e., "non-speech"), system 100 can modify the decision made by mixer 104 and channel gating module 110 to gating open and subsequently gating close such channels, so that the erroneous noise is ultimately not included in the mixed audio signal output from mixer 104. Specifically, in step 212, if erroneous noise is determined to exist in the channel in step 210, then process 200 can continue to step 220. In step 220, due to the detection of erroneous noise, the decision made by mixer 104 and channel gating module 110 to gating open the channel can be modified, and the channel can be gated closed. Process 200 can continue to step 218, where mixer 104 can output a mixed audio signal that does not contain contributions from channels that are now gating closed. In an embodiment, the confidence score from VAD 108 can be used to determine whether the mixer 104's decision to gate open channels can be altered to gate closed channels and / or to affect the relative selection of mixing for each channel in the automatic mixer.
[0028] However, in step 212, if it is determined in step 210 that speech (i.e., "non-noise") is present in the channel, then process 200 may continue to step 214. In step 214, the filter applied in step 208 may be removed, as described in more detail below. In step 216, the channel may be gated open by mixer 104, and in step 218, mixer 104 may output a mixed audio signal containing this channel.
[0029] In this embodiment, steps 210 and 212, used by VAD 108 to identify the presence of speech or noise in a channel, can be performed in parallel or only after mixer 104 and channel gate module 110 have determined a channel gate decision in steps 204 and 206. For example, VAD 108 can collect and buffer audio data from the input audio signal over a predetermined time period to have sufficient information to determine whether the channel contains speech or noise. Thus, during the time period between the decision of mixer 104 and the decision of VAD 108 (regarding whether to modify mixer 104 and channel gate module 110), erroneous noise can temporarily contribute to the mixed audio signal. This contribution of erroneous noise over a short time period can be referred to as front-end noise leakage (FENL). Compared to front-end clipping, FENL occurring in the mixed audio signal can be considered more desirable and less noticeable to listeners of the mixed audio signal. The subjective impact of allowing FENL can be minimized by controlling the amplitude and frequency content of the FENL time period and the selected time length allowed by the FENL.
[0030] In an embodiment, mixer 104 may include a gating state machine that controls the final application of channel gating based on decisions made by mixer 104, channel gating module 110, and VAD 108. The state machine may include: (1) an FEC time period controlled by an algorithmic design outside the design of mixer 104 and channel gating module 110, and delaying the gating on time; (2) a specific duration during the FENL time period, wherein mixer 104 and channel gating module 110 have complete control over channel gating; and / or (3) a final time period, wherein a gating indication from VAD 108 can be logically ANDed with a gating indication from mixer 104 and channel gating module 110. The gating state machine may return to its initial condition when the gating indication from mixer 104 and channel gating module 110 returns to the gating-off channel. Figure 3 The diagram shows a depiction of the gate control state machine.
[0031] Various techniques detailed below can be used to minimize the contribution of FENL to the mixed audio signal by minimizing the energy and spectral contribution of erroneous noise that may temporarily leak into a particular channel. Minimizing the contribution of FENL to the mixed audio signal reduces the impact on speech and voice in the mixed audio signal during the period in which FENL may occur. In some embodiments, this FENL minimization technique can be implemented in premixer 106.
[0032] In some embodiments, premixer 106 may receive status information from voice activity detector 108. The status information may include a combination of an automatic mixer gate flag, a VAD / NAD indicator, and a FENL time period. Premixer 106 may use the status information to determine amplitude attenuation and frequency filtering applied over time. Mixer 104 may receive processed audio signals from premixer 106. The number of processed audio signals from premixer 106 to mixer 104 may be the same as the number of microphones 102 in some embodiments or less than the number of microphones 102 in other embodiments.
[0033] One technique may involve applying attenuation to the gate-on amplitude until VAD 108 can definitively confirm the gate-on channel decision made by mixer 104. Channel attenuation during the FENL time period reduces the impact of error noise while having a relatively insignificant effect on the intelligibility of speech in the mixed audio signal. This technique can be implemented in premixer 106 by applying simple attenuation in step 209 to the channel that the automatic mixer has recently gated within the FENL time window and removing the attenuation application in step 215. The FENL time window exits after a timer expires, corresponding to the length of time during which noise leakage is allowed without materially affecting the subjective audio quality of the speech.
[0034] Another technique could involve reducing the audio bandwidth during the FENL time period. In this case, reducing the audio bandwidth maintains the most important frequencies for the intelligibility of speech or voice in the mixed audio signal during the FENL time period, while significantly reducing the impact of full-band FENL over a period of time (e.g., milliseconds). This technique can be implemented in the premixer 106 by applying a non-voice deemphasis filter in step 208 and removing the application of the non-voice deemphasis filter in step 214, as described above. For example, in step 208, a low-pass filter can be applied after the mixer 104 has made a decision on whether to gate or close the channel (e.g., in steps 204 and 206) but before the VAD 108 makes a decision on the presence of speech or noise in the channel. Once the VAD 108 has made a decision on the presence of speech in the channel (e.g., in steps 210 and 212), the application of the non-voice deemphasis filter can be removed in step 214. In one embodiment, the non-voice deemphasis filter in premixer 106 may be a static second-order Butterworth filter that cross-fades with the unprocessed audio signal from microphone 102. In other embodiments, the non-voice deemphasis filter in premixer 106 may be implemented as two first-order low-pass filters in series, wherein more or less filtering can be applied by shifting the position of the filter poles over time, providing control over the bandwidth of low and high frequencies independently and adaptively over time. This adaptive control of the filters may correspond to FENL timer parameters or VAD confidence metrics. In other embodiments, the non-voice deemphasis filter in premixer 106 may be implemented as a more complex bandwidth-limiting filter that preserves the formant structure of the voice by employing linear predictive coding.
[0035] Another technique may involve altering the crest factor of the audio to minimize perceived noise. Many types of error noise can have crest factors higher than human speech. Sustained high crest factors can be perceived as loudness by humans. By compressing the crest factor of the audio to be equal to or lower than that of human speech during the FENL region, the intelligibility of human speech can be maintained while reducing the perceived loudness of the error noise. In some embodiments, signals with instantaneous temporal crest factors higher than the target can be dynamically compressed to maintain the desired crest factor. In other embodiments, the compression may be modified into a limiter to further ensure that the resulting audio has the desired crest factor.
[0036] Further techniques may include introducing a predetermined amount of FEC, which can psychoacoustically minimize abrupt transient error noise (e.g., pen clicks, book drops on a table, etc.) without significantly affecting the subjective quality of speech (which typically does not exhibit an instantaneous start). In this case, the introduction of FEC can be further refined to simulate the inverse envelope of the transient error noise, which can significantly reduce noise perception without completely removing the speech start that would otherwise be statically decayed during the FENL time period. This can be implemented in step 209 and removed in step 215 by applying time-varying rather than static decay. By using one or more of these techniques, the impact of error noise leakage into the undetected mixed audio signal can be minimized until VAD 108 can make a decision about the presence of speech or noise in the channel. Therefore, this can provide benefits to speech intelligibility without increasing audio path delay.
[0037] The FENL minimization technique described above can be enhanced by using adaptive techniques that automatically modify behavior to better match the operating environment of system 100. Such adaptive techniques can control the timing parameters of the gated state machine described above, as well as parameters such as the inverse FEC envelope shape, bandwidth reduction, attenuation during the FENL time period, FENL minimization time in / out behavior, and / or the time trajectory of mixer 104, to gate channels identified by VAD 108 as containing erroneous noise.
[0038] In an embodiment, system 100 may collect statistics for each channel (corresponding to each of the plurality of microphones 102 (or array lobes)) to identify whether a particular channel contains, on average, speech / voice or noise. For example, in a particular environment, one channel may point to a door, while another channel points to the main location. In this environment, over time, system 100 may determine that the channel pointing to the door is almost exclusively erroneous noise and the channel pointing to the main location is almost exclusively speech. In response, system 100 may tune the channel pointing to the door to apply a longer forced FEC, use more aggressive FENL minimization parameters, and / or cause the gate state machine to give VAD 108 additional priority in gate control decisions. Conversely, system 100 may tune the channel pointing to the main location to eliminate FEC, reduce the use of FENL minimization techniques, and / or cause the gate state machine to provide gate control for mixer 104 for a longer period of time (which in turn forces VAD 108 to have more confidence in its noise-related decisions before changing and gate-closing channels).
[0039] Another technique may include system 100 allowing training and adaptation only when VAD 108 has reached a high confidence threshold level on a specific channel. This can mitigate false positives and / or false negatives in adaptation behavior such as those applied to FENL minimization techniques. A further technique may include system 100 sampling and analyzing audio envelope data of gated open channels subsequently labeled as noise by VAD 108 within an audio cycle in order to update the inverse FEC envelope shape described above.
[0040] In embodiments, adaptive behavior can also be applied to the gating channel shutdown process. For example, during normal voice communication, system 100 may apply a slow ramp to gating the channel shut down to minimize the perception of increases, decreases, or changes in the audio noise floor. As another example, in the presence of noise, system 100 may apply a fast ramp to gating the channel shut down to maximize the effectiveness of gating the channel shut down in response to the decision made by VAD 108. In embodiments, system 100 may combine information from mixer 104 and VAD 108 to determine the reason for gating the channel shut down. This information can be used to dynamically change the speed of channel gating shutdown. Additionally, the non-uniform ramp slope can be used to perceptually optimize both error noise and voice conditions.
[0041] System 100 may include further techniques to address imperfect audio selectivity between microphones 102 (or lobes), which can result in many or all channels having both speech and error noise. In this situation, simply gating off a specific channel containing the highest amount of error noise may not completely eliminate the error noise from the mixed audio signal. This can result in some error noise remaining in the gated channels containing speech. One technique to address this situation may include using a noise leakage filter in the premixer 106. The noise leakage filter can be applied during a period after VAD 108 has made a decision about the presence of speech in a specific channel. If it has been determined that different channels contain error noise (i.e., the decision of mixer 104 to gate the different channels has been changed by VAD 108), then the noise leakage filter can be applied to the channels with speech to mitigate the high-frequency leakage of noise into the channels with speech. In other words, the noise leakage filter can be applied when at least one channel is identified as containing error noise while other channels are identified as not having error noise (i.e., having speech). In one embodiment, the noise leakage filter in premixer 106 may be a static second-order Butterworth filter that cross-fades with the unprocessed audio signal from microphone 102. In other embodiments, the noise leakage filter in premixer 106 may be implemented as two first-order low-pass filters in series, wherein more or less filtering can be applied by shifting the position of the filter poles over time, providing control over the bandwidth of low and high frequencies independently and adaptively over time. This adaptive control of the filters may correspond to the number of other channels identified as noise or a VAD confidence metric. In other embodiments, the noise leakage filter in premixer 106 may be implemented as a more complex bandwidth-limiting filter that preserves the formant structure of the speech by employing linear predictive coding.
[0042] For example, when a specific channel is typically gated off by mixer 104, mixer 104 can attenuate the audio signal in said channel (e.g., by applying -15dB attenuation) to preserve room presence, maintain consistent noise floor as various channels are gated on and off, and reduce the impact of FEC on channels that are later gated on. By using the noise leakage filter described above, system 100 can reduce the bandwidth of gated channels, thus preserving frequencies of voice intelligibility while rejecting frequencies of error noise. This can result in reduced error noise leakage into gated channels.
[0043] In some embodiments, to further reduce the contribution of error noise, when one or more channels are identified by VAD 108 as containing error noise, system 100 may apply additional attenuation (i.e., from -15dB to -25dB) to all gated-off channels and reduce the bandwidth of those channels.
[0044] It should be noted that standard static noise reduction techniques can be used in system 100. In an embodiment, VAD 108 may utilize an audio signal from microphone 102 that has not yet had its noise reduced. VAD 108 uses a non-noise-reduced audio signal, which allows VAD 108 to make its decisions based on the original noise floor of the audio signal, which may be better.
[0045] In this application, the use of transition words is intended to encompass conjunctions. The use of definite or indefinite articles is not intended to indicate cardinality. Specifically, references to “the” object or “a” and “an” object are also intended to indicate one of a possible plurality of such objects. Furthermore, the conjunction “or” can be used to convey simultaneous features rather than mutually exclusive alternatives. In other words, the conjunction “or” should be understood as including “and / or”. The terms “includes,” “including,” and “include” are inclusive and have the same scope as “comprises,” “comprising,” and “comprise,” respectively.
[0046] Any process description or box in the figure should be understood as representing a module, fragment, or portion of code containing one or more executable instructions for implementing a specific logical function or step in the process, and as those skilled in the art will understand, alternative implementations are included within the scope of the embodiments of the invention, wherein, depending on the functionality involved, the functions may be performed in a different order than shown or discussed (including substantially simultaneously or in reverse order).
[0047] This disclosure is intended to illustrate how various embodiments can be modified and used according to the present technology, and not to limit the true, intended, and fair scope and spirit of the invention. The foregoing description is not intended to be exhaustive or limited to the precise forms disclosed. Modifications or variations are possible in view of the teachings above. Several embodiments were chosen and described to provide the best illustration of the principles of the described technology and its practical application, and to enable those skilled in the art to utilize the technology in various embodiments and with various modifications suitable for the particular intended use. All such modifications and variations, when interpreted according to the breadth of their fair, lawful, and impartial authorization, fall within the scope of the embodiments as defined by the appended claims and all their equivalents, as amended during the pending period of this patent application.
Claims
1. A method for detecting speech and noise activity, comprising: Use a mixer to determine the initial gate opening for channels with audio signals; Determine whether there is non-voice audio in the audio signal of the channel initially gated by the mixer, wherein the mixer generates a mixed audio signal based on at least the audio signal of the channel initially gated. During the duration between (1) when the mixer determines whether the gated channel is initially gated and (2) when it determines whether there is non-voice audio in the audio signal of the initially gated channel, one or more of filtering or attenuation is used to minimize erroneous noise in the audio signal of the initially gated channel; and When it is determined that the non-voice audio is present in the audio signal of the channel initially gated open, the mixer is modified by gated close of the channel initially gated open so that the mixer generates the mixed audio signal that does not have the audio signal of the channel initially gated open.
2. The method of claim 1, further comprising applying a non-voice deemphasis filter to the audio signal of the channel initially gated.
3. The method according to claim 2, further comprising: Determine whether there is voice audio in the audio signal of the channel when the gate was initially opened; and When it is determined that the voice audio is present in the audio signal of the channel initially gated, the non-voice deemphasis filter is removed from the audio signal of the channel initially gated.
4. The method of claim 2, further comprising removing the non-voice deemphasis filter from the audio signal of the initially gated channel after the duration between (1) the mixer determining that the gated channel is initially gated and (2) the audio signal of the initially gated channel is present has elapsed.
5. The method of claim 1, further comprising attenuating the audio signal of the channel initially gated.
6. The method of claim 5, further comprising: Determine whether there is voice audio in the audio signal of the channel when the gate was initially opened; and When it is determined that the voice audio is present in the audio signal of the channel initially gated, the attenuation is removed from the audio signal of the channel initially gated.
7. The method of claim 5, further comprising removing the attenuation from the audio signal of the initially gated channel after the duration between (1) the mixer determining that the gated channel is initially gated and (2) the audio signal of the initially gated channel is present has elapsed.
8. The method of claim 1, further comprising applying time-varying attenuation to the audio signal of the channel initially gated.
9. The method of claim 8, further comprising: Determine whether there is voice audio in the audio signal of the channel when the gate was initially opened; and When it is determined that the voice audio is present in the audio signal of the channel initially gated, the time-varying attenuation is removed from the audio signal of the channel initially gated.
10. The method of claim 8, further comprising removing the time-varying attenuation from the audio signal of the initially gated channel after (1) the mixer determines that the duration between the time interval between the initial gated channel being gated and (2) the audio signal of the initially gated channel being gated has elapsed.
11. The method of claim 1, further comprising applying one or more of a crest factor compressor or crest factor limiter to the audio signal of the channel initially gated.
12. The method of claim 11, further comprising: Determine whether there is voice audio in the audio signal of the channel when the gate was initially opened; and When it is determined that the voice audio is present in the audio signal of the channel initially gated, one or more of the crest factor compressor or the crest factor limiter are removed from the audio signal of the channel initially gated.
13. The method of claim 11, further comprising removing one or more of the crest factor compressor or the crest factor limiter from the audio signal of the initially gated channel after the duration between (1) the mixer determining that the gate is initially gated on the channel and (2) determining whether the non-voice audio is present in the audio signal of the initially gated channel has elapsed.
14. The method of claim 1, further comprising, when it is determined that the non-voice audio is present in the audio signal of the channel initially gated, applying additional attenuation to the channel initially gated after the gate is closed.
15. The method of claim 1, further comprising modifying parameters related to minimizing the erroneous noise leakage by using one or more of the filtering or attenuation, based on whether the channel where the initial gate was opened historically contained the non-voice audio or voice audio.
16. The method of claim 1, wherein altering the mixer comprises altering the mixer by controlling the rate at which the gating turns off the channel initially opened by the gating.
17. The method of claim 1, further comprising: Determine whether there is voice audio in the audio signal of the channel when the gate was initially opened; Determine whether there is non-voice audio in the second audio signal of the second channel initially gated by the mixer; and When it is determined that the voice audio is present in the audio signal of the channel initially gated and when it is determined that the non-voice audio is present in the second audio signal of the second channel initially gated, a noise leakage filter is applied to the audio signal of the channel initially gated.
18. The method of claim 1, further comprising determining whether the channel initially gated by the mixer is gated based on (1) a channel selection rule or (2) whether the audio signal of the channel initially gated contains voice audio.
19. A system for detecting speech and noise activity, comprising: A mixer, used to determine the initial gate opening for a channel with an audio signal; An activity detector is configured to determine whether non-voice audio is present in the audio signal of the channel initially gated open by the mixer, wherein the mixer is configured to generate a mixed audio signal based on at least the audio signal of the channel initially gated open. A premixer, which communicates with the mixer, is configured to minimize erroneous noise in the audio signal of the initially gated channel by using one or more of filtering or attenuation during the duration between (1) when the mixer determines that the gate is initially gated on the channel and (2) when the activity detector determines that there is non-voice audio in the audio signal of the initially gated channel. and A channel gate module, which communicates with the activity detector, is configured to modify the mixer to: When the activity detector determines that non-voice audio is present in the audio signal of the channel initially gated, the mixer: The gate closes the channel that was initially opened by the gate; and The mixed audio signal is generated that does not have the audio signal of the channel that was initially gated open.
Citation Information
Patent Citations
Microphone actuation control system suitable for teleconference systems
US4658425A
Microphone actuation control system
US5297210A
Audio signal processor
WO2018211806A1