Ambient noise compensation in teleconferencing

By estimating the spectrum of remote speech and local ambient noise, calculating SII, and adjusting the audio system, the problem of ambient noise impact in teleconferences was solved, improving communication clarity and audio quality.

CN121464481APending Publication Date: 2026-02-03DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480046079.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-17
Filing Date
2024-07-01
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing teleconferencing systems are inadequate in handling ambient noise compensation, which affects speech intelligibility. Improved equipment and methods are needed to enhance communication clarity.

Method used

The speech intelligibility index (SII) is calculated by estimating the spectrum of remote speech and local ambient noise through the control system, and the local audio system is adjusted accordingly, including increasing or decreasing the loudspeaker playback volume to maintain the SII within the target range. Noise compensation is performed by combining steady-state noise estimation and neural network methods.

Benefits of technology

It improves speech intelligibility in teleconferences, enhances communication quality, reduces interference from ambient noise, and provides a clearer audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121464481A_ABST
    Figure CN121464481A_ABST
Patent Text Reader

Abstract

A method of compensating for ambient noise during a teleconference may involve estimating, by a control system, a current speech spectrum corresponding to a speech of a remote teleconference participant; estimating, by the control system, a current noise spectrum corresponding to ambient noise in a local environment in which a local teleconference participant is located; calculating, by the control system, a current speech intelligibility index (SII) based at least in part on the current speech spectrum and the current noise spectrum; determining, by the control system and based at least in part on the current SII, whether to adjust a local audio system used by the local teleconference participant, where the determining involves evaluating the current SII according to one or more target SII parameters; and updating at least one of the one or more target SII parameters in response to a user input corresponding to the playback volume change.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 512,424, filed July 7, 2023, and U.S. Provisional Patent Application No. 63 / 635,570, filed April 17, 2024, all of which are incorporated herein by reference in their entirety. Technical Field

[0002] This disclosure relates to apparatus, systems, and methods for environmental noise compensation (ENC), particularly in the context of teleconference. As used herein, the term "teleconference" encompasses both audio / video teleconference and audio teleconference. Background Technology

[0003] Teleconference has become an integral part of modern life. The ability to communicate clearly during teleconferences relies primarily on speech intelligibility, which in turn depends in part on the presence of noise in the audio signal(s). While existing devices, systems, and methods for speech intelligibility estimation, noise estimation, and ENC offer benefits, improved systems and methods are still desired. Summary of the Invention

[0004] At least some aspects of this disclosure can be implemented via methods. Some such methods involve compensating for ambient noise during a teleconference. For example, some methods may involve estimating, by a control system, a current speech spectrum corresponding to the speech of a remote teleconference participant, and a current noise spectrum corresponding to ambient noise in the local environment of a local teleconference participant. Some methods may involve calculating a current speech intelligibility index (SII) by a control system based at least in part on the current speech spectrum and the current noise spectrum. Some methods may involve determining, by a control system and at least in part on the current SII, whether to adjust the local audio system used by the local teleconference participant. Some methods may involve updating at least one of one or more target SII parameters by a control system in response to user input corresponding to a change in playback volume.

[0005] The determination may involve evaluating the current SII based on one or more target SII parameters. In some examples, the determination may involve determining whether the current SII is within the target SII range. Some methods may involve adjusting at least a portion of the local audio system by a control system in response to determining that an adjustment should be made.

[0006] Some methods may involve determining a confidence value corresponding to the current noise spectrum. The confidence value may, for example, indicate the probability that the current input audio frame primarily corresponds to ambient noise. According to some examples, the confidence value may be a wideband confidence value. Determining whether to adjust the local audio system may be based at least in part on the wideband confidence value. Some methods may involve determining a band-based confidence value for each of a plurality of frequency bands. In some examples, determining whether to adjust the local audio system may involve determining whether to update one or more noise statistics for each frequency band, based at least in part on the band-based confidence value. Some methods may involve determining whether to update one or more noise statistics, based at least in part on the confidence value.

[0007] In some examples, estimating the current noise spectrum may involve estimating the echo coupling gain corresponding to the local loudspeaker playback captured by the local microphone system. In some such examples, estimating the echo coupling gain may involve determining the maximum band loudspeaker reference power for each band of the current audio frame and the previous N-1 audio frames, where N is an integer corresponding to the number of audio frames in the delay line. In some such examples, estimating the echo coupling gain may involve tracking the minimum power of each band of the input microphone signal, and estimating the echo coupling gain for each band of the input microphone signal based at least in part on the maximum band loudspeaker reference power and the minimum power. Some methods may involve determining whether to update the estimated echo coupling gain based at least in part on whether the change in the currently estimated echo coupling gain value relative to the most recently estimated echo coupling gain value exceeds a threshold amount, based on whether the estimated echo coupling gain is updated within a threshold time interval, or both.

[0008] Some methods may involve adjusting at least a portion of the local audio system to maintain the current SII within a target SII range in response to determining that an adjustment should be made. Some methods may involve adjusting at least a portion of the local audio system to maintain the current SII within a target SII range in response to determining that an adjustment should be made. According to some examples, maintaining the current SII within the target SII range may involve increasing the amplifier playback volume of one or more audio frames until the current SII is greater than a first target SII and less than a higher target SII. In some examples, the first target SII may be greater than an intermediate target SII and less than a higher target SII. In some examples, maintaining the current SII within the target SII range may involve decreasing the amplifier playback volume of one or more audio frames until the current SII is less than a second target SII and greater than a lower target SII. According to some examples, the second target SII may be less than an intermediate target SII and greater than a lower target SII.

[0009] Some or all of the operations, functions, and / or methods described herein can be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. Therefore, some innovative aspects of the subject matter described herein can be implemented in non-transitory media on which software is stored.

[0010] At least some aspects of this disclosure can be implemented via an apparatus. For example, one or more devices may be able to perform at least partially the methods disclosed herein. In some embodiments, the apparatus may include an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof. In some examples, the apparatus may be one of the audio devices referenced above. However, in some embodiments, the apparatus may be another type of device, such as a mobile device, a laptop computer, a server, etc.

[0011] Details of one or more embodiments of the subject matter described in this specification are set forth in the following figures and description. Other features, aspects, and advantages will become apparent from the description, figures, and claims. Note that the relative dimensions in the following figures may not be drawn to scale. Attached Figure Description

[0012] Figure 1A An example of a sound source that can be captured by a local microphone during a conference call is shown.

[0013] Figure 1B This is a block diagram illustrating examples of components of an apparatus capable of implementing various aspects of this disclosure.

[0014] Figure 1C It is a flowchart outlining examples of methods that can be performed by means or systems such as those disclosed herein.

[0015] Figure 2A An example box of a novel noise estimator based on some examples is shown.

[0016] Figure 2B The following are embodiments according to some disclosed implementations. Figure 2A The box of the main noise estimator.

[0017] Figure 3A The following are embodiments according to some disclosed implementations. Figure 2B The box for the echo coupling gain estimator.

[0018] Figure 3B A graph showing the SII estimate over time based on an example is shown.

[0019] Figure 4 A graph showing the SII over time as estimated based on another example is shown.

[0020] Figure 5 Another graph showing the estimated SII over time is shown.

[0021] Figure 6 This is a flowchart outlining another example of a method that can be performed by an apparatus or system such as those disclosed herein.

[0022] In the various figures, similar reference numerals and names indicate similar elements. Detailed Implementation

[0023] Figure 1A Examples of sound sources that can be captured by a local microphone during a conference call are shown. In these examples, the local speech 102 of the local conference call participants, ambient noise 104 (also referred to herein as “background noise” or “surround noise”), and amplifier playback sound 106 (also referred to herein as “echo”) from amplifier 108 are captured by microphone 112. During the conference call, the speech of remote conference call participants (also referred to herein as “remote speech”) is played back by amplifier 108. The term “remote conference call participant” refers to a conference call participant located at a location other than that of the local conference call participants.

[0024] Figure 1A An example of a noise estimator 114 is also shown, and examples of said noise estimator are disclosed herein. The noise estimator 114 can be, for example, via a reference... Figure 1B An example of the control system 110 described is used for implementation.

[0025] The speech of remote conference call participants and the local speech 102 of local conference call participants are typically time-divided. Instances of “double-speech” where the remote speech and local speech 102 overlap in time are atypical. In most cases, double-speech occurs only when a conference call participant wants to interrupt. To prevent recaptured remote speech from being transmitted back to the remote end, echo management is typically employed in the signal chain, and the echo signal 106 is usually significantly suppressed.

[0026] Reliable noise estimation and denoising techniques can help conference participants better understand what other participants in a teleconference are saying. This paper discloses various improved noise estimation and denoising techniques.

[0027] Figure 1B This is a block diagram illustrating examples of components of an apparatus capable of implementing various aspects of this disclosure. As with the other figures provided herein, Figure 1B The types and quantities of elements shown are provided as examples only. Other embodiments may include more elements, fewer elements, elements of different types, or combinations thereof. According to some examples, device 101 may be or may include a device configured to perform at least some of the methods disclosed herein, such as a smart audio device, laptop computer, cellular phone, tablet device, smart home hub, etc. In some such embodiments, device 101 may be or may include a server configured to perform at least some of the methods disclosed herein.

[0028] In this example, apparatus 101 includes interface system 105 and control system 110. In some embodiments, control system 110 may be configured to perform at least partially the methods disclosed herein. In some embodiments, control system 110 may be configured to compensate for ambient noise during a teleconference.

[0029] In some examples, the control system 110 may be configured to estimate a current speech spectrum corresponding to the speech of a remote teleconference participant. According to some examples, the control system 110 may be configured to estimate a current noise spectrum corresponding to ambient noise in the local environment where a local teleconference participant is located. In some examples, the control system 110 may be configured to calculate a current speech intelligibility index (SII) based at least in part on the current speech spectrum and the current noise spectrum. In some examples, the control system 110 may be configured to determine, at least in part on the current SII, whether to adjust the local audio system used by the local teleconference participant. In some examples, the determination may involve evaluating the current SII according to one or more target SII parameters. According to some examples, the control system 110 may be configured to update at least one of one or more target SII parameters in response to user input corresponding to a change in playback volume.

[0030] Interface system 105 may include one or more network interfaces and / or one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). According to some embodiments, interface system 105 may include one or more wireless interfaces. Interface system 105 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, interface system 105 may include a control system 110 and a memory system (such as...) Figure 1BOne or more interfaces between the optional memory system 115 shown. However, in some instances, the control system 110 may include a memory system.

[0031] For example, the control system 110 may include a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.

[0032] In some implementations, the control system 110 may reside in more than one device. For example, a portion of the control system 110 may reside in a device within the environment (such as a laptop computer, tablet computer, smart audio device, etc.), and another portion of the control system 110 may reside in a device outside the environment (such as a server). In other examples, a portion of the control system 110 may reside in a device within the environment, and another portion of the control system 110 may reside in one or more other devices within the environment.

[0033] Some or all of the methods described herein can be executed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may, for example, reside on... Figure 1B In the optional memory system 115 and / or control system 110 shown. Therefore, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media on which software is stored. For example, the software may include instructions for controlling at least one device to process audio data. For example, the software may be provided by, for example, Figure 1B The control system 110 and other control system components perform the operation.

[0034] In some examples, device 101 may include Figure 1B The optional microphone system 120 is shown. The optional microphone system 120 may include one or more microphones. In some embodiments, the one or more microphones may be part of or associated with another device, such as a loudspeaker, smart audio device, etc.

[0035] According to some embodiments, device 101 may include Figure 1BThe optional loudspeaker system 125 is shown. The optional loudspeaker system 125 may include one or more loudspeakers. Loudspeakers may sometimes be referred to herein as “speakers”. In some examples, at least some of the loudspeakers of the optional loudspeaker system 125 may be arbitrarily positioned. For example, at least some of the speakers of the optional loudspeaker system 125 may be placed in locations that do not correspond to any standard-defined speaker layout, such as Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4, Dolby 9.1, Hamasaki 22.2, etc. In some such examples, at least some of the loudspeakers of the optional loudspeaker system 125 may be placed in locations convenient for space (e.g., where there is space to accommodate the loudspeakers), but not in any standard-defined loudspeaker layout.

[0036] In some embodiments, device 101 may include Figure 1B The optional sensor system 130 is shown. The optional sensor system 130 may include a touch sensor system, a gesture sensor system, one or more cameras, etc.

[0037] In some embodiments, device 101 may include Figure 1B The optional display system 135 is shown. The optional display system 135 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some instances, the optional display system 135 may include one or more organic light-emitting diode (OLED) displays. In some examples where device 101 includes display system 135, sensor system 130 may include a touch sensor system and / or gesture sensor system for proximity to one or more displays of display system 135. According to some such embodiments, control system 110 may be configured to control display system 135 to present a graphical user interface (GUI), such as a GUI associated with implementing one of the methods disclosed herein.

[0038] As mentioned above, reliable noise estimation and denoising techniques can help conference participants better understand what other participants in a teleconference are saying. Recent developments involving neural network-based methods have contributed to solving the denoising problem. Neural network-based methods are generally able to remove many types of noise, provided that the noise during the neural network training process has provided the neural network with sufficient exposure to each specific type of noise to be removed. However, training a neural network to remove every possible type of noise can be difficult or even impossible.

[0039] Therefore, noise compensation remains an important aspect of providing acceptable audio for teleconferences. The basic principle of noise compensation is to adjust the playback volume based on ambient noise sensed by the microphone. In some examples, noise compensation can be applied on a per-band basis, where different gains are applied to different bands.

[0040] In the context of noise compensation, robust and stable steady-state noise estimators are crucial. The quality of the steady-state noise estimator typically determines the overall system performance. The basic functions of a steady-state noise estimator include: • Steady-state noise estimators should only estimate “steady-state” noise to prevent rapid volume adjustments in the face of incidental (dynamic) noise (e.g., dog barking, sounds from tapping on a table, sounds from keyboard typing, etc.). As used herein, the term “steady-state noise” does not refer to a type of noise with “strict stationarity” where its statistical properties do not change over time, but rather to a type of noise with “broad stationarity” where the first moment and autocovariance do not change over time and where the second moment is finite for all times. Steady-state noise is a type of “steady-state process” that is a stochastic process whose unconditional joint probability distribution does not change over time.

[0041] • The steady-state noise estimator should only estimate the "real" ambient noise of the local environment, not the noise in the replayed echo signal. For example, refer to Figure 1A The steady-state noise estimator should only estimate ambient noise 104, not loudspeaker playback or "echo" sound 106. If the estimated noise is actually mainly caused by echo (echo noise is dominant), the ambient noise compensation (ENC system) will form a positive feedback loop, which will cause any gain adjustment to produce a maximum or minimum volume setting.

[0042] Figure 1C This is a flowchart outlining examples of methods that can be performed by means or systems such as those disclosed herein. As with other methods described herein, the blocks of method 140 need not be performed in the indicated order. In some embodiments, one or more blocks of method 140 may be performed simultaneously. Furthermore, some embodiments of method 140 may include more or fewer blocks than those shown and / or described.

[0043] According to this example, method 140 is a method for compensating for ambient noise during a teleconference. The block of method 140 can be performed by one or more devices, which can be (or may include) a control system, such as... Figure 1B The control system 110 is shown and described above. According to this example, the frame of method 140 is repeated for each new audio data frame.

[0044] In this example, method 140 involves calculating the current Speech Intelligibility Index (SII) and, if necessary, making incremental changes in the system (such as changes in playback volume) to keep the system within the Speech Intelligibility Index target. The system may include a device configured to provide teleconferences. In some examples, the system may include a laptop computer with a speaker and microphone, configured to provide teleconferences and used by local participants during the teleconference. The output of the speaker during the teleconference is an example of what is referred to in this document as “the speech of a remote teleconference participant.” See again Figure 1A The audio captured by the local microphone will include ambient noise 104 and speech 102 from local conference call participants. For a noise detector, speech 102 may be considered a type of interference.

[0045] according to Figure 1C In the example shown, box 145 relates to obtaining an audio data frame from a local microphone. In this example, box 150 relates to estimating the speech spectrum corresponding to speech 102 in the current audio data frame from the microphone. According to this example, box 155 relates to calculating the current noise spectrum corresponding to the ambient noise 104 in the current audio data frame from the microphone. In some examples, boxes 150 and 155 may be performed simultaneously. According to some examples, the noise spectrum is based in part on historical data. In some examples, box 155 will keep the noise spectrum updated as new data frames are received. In some examples, the current noise spectrum determined by box 155 is for the current audio data frame from the microphone and is used to determine the SII and other parameters. In this example, box 160 relates to calculating the current SII, which is the SII corresponding to the current audio data frame from the microphone. According to this example, box 165 relates to determining whether to adjust the local audio system used by local teleconference participants, based at least in part on the current SII.

[0046] In some examples, box 150 may involve estimating the speech spectrum, as in the original publication in 1969 and revised by the American National Standards Institute and the Acoustical Society of America in 1997. Methods for Calculation of the Speech Intelligibility IndexThe method described in [Methods for Calculating Speech Intelligibility Indices] (hereinafter referred to as ASA / ANSI S3.5-1997) is hereby incorporated by reference. For example, box 150 may relate to estimating the speech spectrum as described in the section “Methods for determining input variables for SII calculation” on pages 11–13 of ASA / ANSI S3.5-1997. However, in some alternative examples, box 150 may relate to estimating the speech spectrum according to one or more other methods.

[0047] In some alternative examples, box 150 may involve estimating the speech spectrum according to a modified version of what is described in ASA / ANSI S3.5-1997. Some such examples involve using “insertion gain” which differs in some respects from what is described in ASA / ANSI S3.5-1997. In ASA / ANSI S3.5-1997, for an amplifying or attenuating device worn by a listener, the insertion gain at a given frequency is the difference (in decibels) between the pure-tone sound pressure level (SPL) at the eardrum when the amplifying / attenuating device is in place and the pure-tone SPL at the eardrum when the amplifying / attenuating device is removed. The insertion gain is used to calculate the equivalent speech spectrum level (see Section 5.1.3, starting on page 12 of ASA / ANSI S3.5-1997) and also to calculate the equivalent noise spectrum level (see Section 5.1.4, page 13 of ASA / ANSI S3.5-1997). For example, according to ASA / ANSIS 3.5-1997, if the speech spectrum is measured at the center of the listener's head, the equivalent speech spectrum for a particular frequency band is calculated as the measured speech spectrum level for that frequency band plus the insertion gain for that frequency band.

[0048] Some publicly available examples involve mapping the equivalent SPL of the playback to this insertion gain. In some examples, the current hardware / software gain applied to the amplifier is converted to the equivalent SPL at the listener's location. According to some examples, this conversion is assisted by tuning parameters during the tuning process of a particular device (such as a particular laptop model). In some examples, this equivalent SPL is subtracted from the nominal SPL associated with the "normal" speech spectrum. The final subtracted value is used as the insertion gain, where the same value is used for all frequency bands.

[0049] In some examples, box 155 may involve calculating the current noise spectrum corresponding to the ambient noise 104 in the current audio data frame from the microphone, according to a method described in ASA / ANSI S3.5-1997 (e.g., on page 13). However, in some disclosed examples, box 155 may involve calculating the current noise spectrum according to an alternative method.

[0050] According to some such methods, box 155 involves calculating the current noise spectrum corresponding to the novel noise estimator (NE), referring to... Figures 2A to 3A Examples of the novel NE are described in more detail. Some implementations of the new noise estimator include at least a first noise estimator or a primary noise estimator and a second auxiliary noise estimator. The second noise estimator can be configured to respond to changes in noise level relatively faster than the first noise estimator or the primary noise estimator.

[0051] In some examples, box 160 may involve determining the current SII as a value between 0 and 1 based on the current speech spectrum estimated by box 150 and the current noise spectrum estimated by box 155, where 1 represents very intelligible and 0 represents completely incomprehensible. In some examples, box 160 may involve determining the current SII based on the current speech spectrum, the current noise spectrum, and a measured or hypothesized hearing threshold level.

[0052] According to some such examples, the SII calculated in box 160 can be the SII metric described in ASA / ANSI S3.5-1997. In some such examples, the SII metric can be calculated as described in the section “Methods for calculating Speech Intelligibility Index” on pages 9-11 of ASA / ANSI S3.5-1997, which is specifically incorporated herein by reference.

[0053] Based on some examples, the current SII can be determined as follows:

[0054] In the aforementioned equation, sii[current] represents the current smoothed SII value, sii[current-1] represents the smoothed SII value calculated in the previous block, Sii_alpha represents the tuning parameter, and Instanteous_sii represents the instantaneous output of the metric described in ASA / ANSI S3.5-1997. In one example, the tuning parameter Sii_alpha could be 2^(-0.02 / 0.15), where 0.02 represents the block length in seconds. Other examples may use larger or smaller tuning parameters.

[0055] Figure 2A An example box of a novel noise estimator based on some examples is shown. In this example, the noise estimator 200 includes a first noise estimator 210 (also referred to herein as the main noise estimator 210), a second noise estimator 220 (also referred to herein as the auxiliary noise estimator 220), and a fast tracking flag module 230. According to this example, the main noise estimator 210, the auxiliary noise estimator 220, and the fast tracking flag module 230 are derived from a reference... Figure 1B An example of the control system 110 described herein is used for implementation. Similar to the other figures provided herein, Figure 2A The types and quantities of elements shown are provided as examples only. Other implementations may include more elements, fewer elements, different types of elements, different arrangements of elements, or combinations thereof.

[0056] exist Figure 2A In the example shown, both the primary noise estimator 210 and the secondary noise estimator 220 are configured to receive loudspeaker audio data 222 and microphone audio data 224, for example, as shown in the reference. Figure 1C As described in box 145. In some examples, the loudspeaker audio data 222 may be or may include a reference signal corresponding to the content played back by the local loudspeaker.

[0057] According to this example, the main noise estimator 210 is configured to determine and output a noise estimate 208 and a confidence metric 207 (also referred to herein as a “confidence value”) based at least in part on the loudspeaker audio data 222 and the microphone audio data 224. In this example, the noise estimate 208 and the confidence metric 207 correspond to the current noise spectrum, which in turn corresponds to the current input audio frame of the microphone audio data 224. The confidence metric 207 indicates that the current input audio frame primarily corresponds to ambient noise (as referenced). Figure 1A The probability of ambient noise 104 described. In some examples, the master noise estimator 210 may be configured to determine the noise estimate 208, the confidence metric 207, or both, at least in part, based on the fast tracking flag 228 from the fast tracking flag module 230. See below for reference. Figure 2B and Figure 3A The example blocks and functions of the main noise estimator 210 are described in more detail.

[0058] In this example, the auxiliary noise estimator 220 is configured to determine and output a noise estimate 226 based at least in part on the loudspeaker audio data 222 and the microphone audio data 224. Here, the noise estimate 226 corresponds to the current input audio frame of the microphone audio data 224. According to some published examples, the auxiliary noise estimator 220 is configured to respond to changes in noise level and generate a response noise estimate 226 relatively faster than the primary noise estimator 210, thereby improving the response speed (reducing response time) of the noise estimator 200.

[0059] According to this example, the fast tracking flag module 230 is configured to determine and output the fast tracking flag 228 based at least in part on noise estimates 226 and 208. In some examples, if the main noise estimator 210 receives the fast tracking flag 228, the main noise estimator 210 can be configured to converge to the new noise level relatively faster.

[0060] In some examples, if the noise estimate 226 from the secondary noise estimator 220 indicates that the power of the broadband noise is greater than the power of the broadband noise indicated by the noise estimate 208 from the primary noise estimator 210 by a certain threshold and this difference persists for a defined time interval, the fast tracking flag module 230 is configured to set the value of the fast tracking flag 228 to its maximum value, for example, to 1. In other words, when it is determined that the difference between (1) the noise power estimated by the secondary noise estimator 220 and (2) the noise power estimated by the primary noise estimator 210 is greater than a difference threshold for at least a first time threshold, the fast tracking flag module 230 can be configured to set the value of the fast tracking flag 228 to its maximum value, for example, to 1. According to some examples, when it is determined that the difference between (1) and (2) is less than a difference threshold for at least a second time threshold, the fast tracking flag module 230 can be configured to set the value of the fast tracking flag 228 to its minimum value, for example, to 0. The first time threshold may or may not be equal to the second time threshold. In some examples, the power threshold can be in the range of 1 dB to 5 dB, for example, 1 dB, 2 dB, 3 dB, 4 dB, or 5 dB. According to some examples, the time constant can be in the range of 0.5 seconds to 2 seconds, for example, 0.5 seconds, 0.6 seconds, 0.7 seconds, 0.8 seconds, 0.9 seconds, 1.0 seconds, 1.1 seconds, 1.2 seconds, 1.3 seconds, 1.4 seconds, 1.5 seconds, 1.6 seconds, 1.7 seconds, 1.8 seconds, 1.9 seconds, or 2 seconds.

[0061] According to some examples, the auxiliary noise estimator 220 can be configured with different tuning parameters than the main noise estimator 210. In some alternative examples, the auxiliary noise estimator 220 can be based on a deep neural network. Such a noise estimator can have a very fast response.

[0062] Figure 2B The following are embodiments according to some disclosed implementations. Figure 2A The main noise estimator is shown in the box. In this example, the main noise estimator 210 includes a speech activity detector (VAD) 201, an echo coupling gain estimator 202, a confidence calculation block 203, and a noise and confidence statistics calculation block 204. According to this example, VAD 201, echo coupling gain estimator 202, confidence calculation block 203, and noise and confidence statistics calculation block 204 are derived from a reference... Figure 1B An example of the control system 110 described herein is used for implementation. Similar to the other figures provided herein, Figure 2B The types and quantities of elements shown are provided as examples only. Other implementations may include more elements, fewer elements, different types of elements, different arrangements of elements, or combinations thereof.

[0063] In this example, VAD 201 is configured to detect local speech activity based on microphone audio data 224 and send a speech activity signal 211 to an echo-coupled gain estimator 202, a confidence calculation block 203, and a noise and confidence statistics calculation block 204. According to some examples, the echo-coupled gain estimator 202, the confidence calculation block 203, and the noise and confidence statistics calculation block 204 only activate when the speech activity signal 211 indicates the absence of local speech activity.

[0064] According to this example, the echo coupling gain estimator 202 is configured to estimate the coupling gain from loudspeaker playback to microphone capture for each of multiple frequency bands based on microphone audio data 224 and loudspeaker audio data 222, and provide the corresponding coupling gain estimate 205 to the confidence calculation block 203. In other words, referencing Figure 1A In the scenario depicted, the echo coupling gain estimator 202 is configured to estimate the contribution of the echo 106 to the sound detected by the microphone 112 and present in the microphone audio data 224, and provide a corresponding coupling gain estimate 205. Figure 2B In the example shown, the loudspeaker audio data 222 is or includes a reference signal corresponding to the content played back by the local loudspeaker. Figure 3A More details of the echo coupling gain estimator 202 based on some examples are shown.

[0065] In this example, confidence calculation block 203 is configured to calculate, for each frequency band, the probability that ambient noise dominates in the current input audio frame of microphone audio data 224. According to this example, confidence calculation block 203 is configured to output confidence data 206 to noise and confidence statistics calculation block 204. In this example, confidence data 206 includes at least a confidence value corresponding to the current noise spectrum. Here, the confidence value indicates the probability that the current input audio frame mainly corresponds to ambient noise. According to this example, confidence data 206 also includes a binary value (e.g., 0 or 1) indicating whether the current input audio frame is likely to be mainly ambient noise and should be processed accordingly.

[0066] According to this example, the noise and confidence statistics calculation block 204 is configured to determine and output a noise estimate 208 and a confidence metric 207 based at least in part on microphone audio data 224 and confidence data 206. The confidence metric 207 indicates that the current input audio frame primarily corresponds to ambient noise (as referenced). Figure 1A The possibility of environmental noise (104) described. In some examples, the noise and confidence statistics calculation block 204 can be configured to be at least partially based on noise from... Figure 2A The fast tracking flag module 230 and the fast tracking flag 228 ( Figure 2B (Not shown in the image) to determine noise estimation 208, confidence metric 207, or both. Detailed examples of how noise and confidence statistics calculation block 204 can work are provided below.

[0067] Figure 3A The following are embodiments according to some disclosed implementations. Figure 2B The box of the echo coupling gain estimator. In this example, the echo coupling gain estimator 202 includes a delay line module 301, a minimum follower 302, a threshold detector 303, a maximum follower 304, an update control block 305, and a subtraction node 312. According to this example, the delay line module 301, the minimum follower 302, the threshold detector 303, the maximum follower 304, the update control block 305, and the subtraction node 312 are defined by a reference... Figure 1B An example of the described control system 110 is used for implementation. According to this example, the echo coupling gain estimator 202 operates only when there is no local speech activity, as determined by VAD 201 (see [reference]). Figure 3A ). As with the other figures provided in this article, Figure 3A The types and quantities of elements shown are provided as examples only. Other implementations may include more elements, fewer elements, different types of elements, different arrangements of elements, or combinations thereof.

[0068] exist Figure 3AIn the example shown, delay line module 301 is a delay line of length N and is configured to receive loudspeaker audio data 222, which is or includes a loudspeaker reference signal corresponding to content played back by a local loudspeaker. In some examples, N can be in the range of 8 to 16 frames, for example, 8, 9, 10, 11, 12, 13, 14, 15, 16 frames, etc. According to some examples, each frame can be in the range of 10 milliseconds (ms) to 30 ms, for example, 10 ms, 12 ms, 14 ms, 16 ms, 18 ms, 20 ms, 22 ms, 24 ms, 26 ms, 28 ms, 30 ms, etc. According to this example, the input loudspeaker reference signal is in the form of a bandgap loudspeaker reference power. In this example, delay line module 301 is configured to perform a maximum operation (MAX) across the entire delay line. The purpose of passing the amplifier reference signal through the delay line is to compensate for the natural delay between the amplifier reference signal and the corresponding microphone capture of the sound reproduced by one or more local amplifiers, so as to take into account sound wave propagation delay, circuit delay, software buffer delay, and reverberation buildup that may be present in the local room. In this example, the delay line module 301 outputs maximum power data 310, which indicates the maximum value of the amplifier reference signal power for each frequency band of the current audio frame and the previous N-1 audio frames.

[0069] According to this example, the minimum power follower 302 is configured to track the minimum power of each input frequency band in each frame of the microphone audio data 224 and output minimum power data 311, which indicates the minimum input power of each frequency band in each frame of the microphone audio data 224. Figure 3A In the text, "B" indicates the number of frequency bands. In some examples, the time window size of the minimum follower 302 can be dynamically adjusted. For example, (see again...) Figure 2A In some implementations, when the fast tracking flag 228 is input to the master noise estimator 210, the minimum follower 302 will shorten its window size to help the master noise estimator 210 converge to the solution more quickly. In one such example, the minimum follower 302 may shorten its window size from 1.0 second to 250 ms. Other examples may involve different initial window sizes, different shortened window sizes, or both.

[0070] exist Figure 3AIn the example shown, threshold detector 303 is a simple threshold detector configured to determine when the power of the current frame of the input microphone audio data 224 is at least a threshold amount higher than a tracked minimum power level, and to provide a corresponding threshold detector output 313. The tracked minimum power level is highly dependent on the sensitivity of a particular microphone. Therefore, the range of tracked minimum power levels may vary from microphone to microphone. According to some examples, the threshold amount is in the range of 3 dB to 10 dB, for example, 3 dB, 4 dB, 5 dB, 6 dB, 7 dB, 8 dB, 9 dB, or 10 dB. In some examples, the threshold detector output 313 of threshold detector 303 is 1 whenever the power of the current frame of the input microphone audio data 224 is at least a threshold amount higher than the tracked minimum. In some such examples, whenever the power of the current frame of the input microphone audio data 224 is at least a threshold amount higher than the tracked minimum, the threshold detector output 313 of threshold detector 303 is 1. No When the value is at least a threshold amount higher than the tracked minimum, the threshold detector outputs a value of 0 for 313.

[0071] In this example, subtraction node 312 is configured to subtract (a) the input power of the current frame of the input microphone audio data 224 from (b) the maximum power of the loudspeaker reference signal, as determined by the maximum power data 310 output by delay line module 301, to produce subtraction node output 315. Subtraction node output 315 represents a potential coupling gain estimate.

[0072] According to this example, the maximum value follower 304 is configured to determine and output the coupling gain estimate 205 based on the subtraction node output 315, and update the control signal 317 from the update control module 305.

[0073] In this example, update control module 305 is configured to determine whether to disallow or allow the maximum follower 304 to output subtraction node output 315 as the current coupling gain estimate 205. According to some such examples, update control signal 317 controls whether the subtraction node output 315 will be received by the maximum follower 304. Although, according to this example, echo coupling gain estimator 202 only operates when there is no local speech activity, local transient noise may still be included in the microphone signal 224. Whenever some transient noise is present in the frequency band, the gain calculated by subtraction node 312 will be inaccurate and should be eliminated.

[0074] However, without prior knowledge of the coupling gain, determining whether the current input audio frame contains local transient noise or is solely generated by a strong echo (in other words, by loud loudspeaker playback) can be challenging. And this coupling gain is what the echo coupling gain estimator 202 needs to estimate here.

[0075] This problem can be overcome by using the fact that the coupling gain typically does not change abruptly and continuously over at least N frames, where N is the delay line length of the delay line module 301. Therefore, according to some examples, whenever a new maximum gain exists, the control system 110 only allows the maximum gain follower 304 to track the new maximum gain if the following two conditions are met: 1. The new maximum gain is within a specific range of the previously tracked maximum gain, such as 1 dB, 2 dB, 3 dB, 4 dB, etc. In some implementations, if no initial value is measured (see below), this condition can be ignored for the first update; 2. No new maximum gain was achieved within the first N frames.

[0076] The aforementioned problems can also be overcome by considering that coupling gain is primarily used for devices participating in teleconferences (such as telephones, laptops, conference endpoints, etc.). While coupling gain may be affected by the acoustic properties of the room where the device is located, it is mainly determined by the industrial design of the device itself. This coupling gain is available once the product is manufactured, and in some examples, it can be stored for future use as an initial value.

[0077] Example of confidence calculation block function This section includes how it can be implemented. Figure 2B An example of confidence calculation block 203. Given the coupling gain of microphone audio data 224 and the current audio frame, confidence calculation block 203 can calculate the probability that the current audio frame is mainly ambient noise. Assume the coupling gain (in dB) of frequency band b is... And the maximum reference power is 310. (In dB). Estimated echo power in dB. It can be represented as follows:

[0078] Microphone audio data 224 current audio frame (power is The probability that a sound is ambient noise can be represented as follows:

[0079] In the aforementioned equation, The threshold value can be 4 dB, 5 dB, 6 dB, 7 dB, 8 dB, etc.

[0080] Based on some examples, confidence calculation block 203 can determine the indicator. (Corresponding to frequency band b) and sends it to the noise and confidence statistics calculation block 204. In some such examples, the confidence calculation block 203 can determine the indicator flag by implementing the following set of conditions. :

[0081] In the aforementioned equation, This represents the estimated noise mean for frequency band b. The second condition (if...) and This allows noise statistics to be updated when the noise disappears. Even if the possibility that the current frame is ambient noise is questionable, the estimated level may be unreasonable under such conditions.

[0082] Noise and confidence level statistical calculations This section includes how it can be implemented. Figure 2B Example of noise and confidence score statistics calculation block 204. Figure 3A In the example shown, the input to the noise and confidence statistics calculation block 204 includes minimum power data 311 from the minimum follower 302. Minimum power data 311 indicates the minimum input power for each frequency band of each frame of the microphone audio data 224. According to this example, the input to the noise and confidence statistics calculation block 204 also includes the microphone signal power. and indicator signs Both of these correspond to frequency band b.

[0083] To prevent transient noise from being controlled and to improve estimation accuracy, in some examples, the noise and confidence score calculation block 204 is only used for... And will only be implemented if the following conditions are met. Included in the cumulative total:

[0084] In the aforementioned equation, Indicates what is being tracked The minimum value, and This represents a threshold, such as 8 dB, 9 dB, 10 dB, 11 dB, 12 dB, etc. In some examples, the confidence statistics are only updated when the noise accumulator is updated. As input, according to some examples, the output of the noise and confidence statistics calculation block 204 includes the estimated ambient noise power for each frequency band. and the confidence value corresponding to the current ambient noise power estimate .

[0085] Possible uses of the output from the noise and confidence statistics calculation block In some implementations, confidence value It can be used to generate broadband confidence labels. P To control the behavior of noise compensation control logic, for example, as follows:

[0086] In the aforementioned equation, Denotes the weighting factor, where, ,and H This indicates that the hysteresis function with two hysteresis curve thresholds outputs either 0 or 1. According to some implementations, the compensation logic or mechanism (multiple) should only be applied when... Time-based activities. In some examples, whenever... And when the current volume (or / other settings) is not in the user's (multiple) settings, (multiple) compensation logic or mechanisms can restore (multiple) settings to the user's (multiple) settings.

[0087] Determine and apply system adjustments After the control system has estimated the current SII, in some implementations, the control system will perform a multi-step process to determine what the target SII is, whether the system needs to be changed, and if so, how to change the system to achieve the target SII. In some examples, this process corresponds to Figure 1C Box 165. Therefore, this section describes various actions that can be performed according to various embodiments of box 165.

[0088] In some examples, box 165 can calculate one or more changes and apply them to the system, such that... Figure 1C The SII calculated in box 160 is within the target SII range. This variation may include, but is not limited to, the following: • Changes in speaker hardware / software gain (e.g., volume control); • An increase in gain will correspond to an increase in SII. The logic is similar for a decrease; • Changes in equalization (EQ); • Changes to the volume leveler effect in Dolby Audio Processing (DAP).

[0089] In some examples, this control(s) will only be implemented when the broadband confidence indicator is high (e.g., 1), in which case there is an estimated noise spectrum that actually comes from a high confidence level of background noise. If the confidence indicator is low (e.g., 0) and none of the above controls are in the user setting, in some examples the control system will restore the user setting(s).

[0090] The following are various examples of how to ensure that SII remains within the target SII range. Figure 3B A graph showing the estimated SII over time based on an example is shown. The estimated SII can, for example, correspond to... Figure 1C The output of frame 160. Figure 3A The diagram illustrates how SII can fluctuate over a period of time. During this period, in this example, box 165 does not recommend any (multiple) system changes, as SII remains within the high and low target ranges, which... Figure 3B The values ​​are shown as "High Target SII" and "Low Target SII". In this case, box 165 is in a state referred to herein as "ref_gain_adj = NONE", indicating that no gain adjustment will be applied to the audio played back on the local system that provides the teleconference to local teleconference participants.

[0091] Figure 4 A graph showing the SII over time as estimated based on another example is shown. Figure 4 This illustrates what happens when SII increases beyond high_target_sii. Between the start time and t1, box 165 does not suggest any volume change (ref_gain_adj = NONE). At t1, the SII value has exceeded the range, surpassing high_target_sii. In this case, box 165 will suggest a system change and will continue to do so in subsequent frames (with optional pauses between suggestions, as described below) until SII is between dec_aim_target_sii and low_target_sii Between. Before this is achieved, box 165 is in the ref_gain_adj = INC_REQ state.

[0092] Based on knowledge about how the system will react and how SII will tend to fluctuate. dec_aim_target_sii Values ​​can be compared with target_sii Same, or can be with target_sii Different. For example, in some systems, it may be known that SII may be underestimated compared to the true SII due to the slow adaptation of noise estimation. In these cases, dec_aim_target_sii, slightly lower than target_sii (but higher than low_target_sii), may be preferred. Once SII falls between dec_aim_target_sii and low_target_sii (at t2), in this example, the system state will return to the reference state. Figure 3B The behavior described. In our implementation, dec_aim_target_sii is the same as target_sii.

[0093] An example of a possible change to the system is reducing the volume. The volume reduction can be specified such that the reduction is proportional to the difference between the current SII and our target SII. In one implementation, the control system can determine possible gain changes in dB, for example, possible gain changes such that:

[0094] In the aforementioned equation, Gain_dec_delta represents the first tuning parameter, which is 0.6 in one example, and Dec_speed represents the second tuning parameter, which is 1.0 in one example. Other examples may involve different tuning parameters. In some alternative examples, the first tuning parameter may be 0.5, 0.55, 0.65, 0.7, etc., and the second tuning parameter may be 0.9, 0.95, etc.

[0095] In some implementations, whenever the control system determines a possible change in the system, it can be configured to optionally wait for a time interval measurable in the input audio frame to allow the system to stabilize before checking the current SII and suggesting a new change. The suggested change may or may not be implemented. For example, in some examples, the suggested change will not be implemented if a feature corresponding to one or more disclosed methods has been turned off or disabled (e.g., based on user input). In some such implementations, the control system's waiting time between checking the current SII and suggesting a new change is configured such that the amount of time spent waiting can be proportional to the time between the SII and the time between the SII and the time between the SII and the time between the SII and the SII. dec_aim_target_sii The degree of proximity is proportional. In some implementations, the control system can first calculate:

[0096] as well as

[0097] Based on the value of Diff_to_target, the control system can calculate wait_factor, for example as follows: Wait_factor = 1 if 0>= diff or diff>= Diff_range Wait_factor = 1, if diff > Diff_range / 3 Wait_factor = 2, if diff > Diff_range / 5 Wait_factor = 3, if diff > Diff_range / 7 Wait_factor = 4, in all other cases.

[0098] In some such examples, the number of input audio frames corresponding to the waiting time interval is equal to:

[0099] In the aforementioned equation, gain_adj_holdon_frames represents the tuning parameters. In one implementation, the tuning factor is 96 for a 20 ms block length. In other implementations, the tuning factor can be 90, 92, 94, 98, or 100 for a 20 ms block length.

[0100] In some instances, the following may occur: • At t0, SII is within high_target_sii and low_target_sii.

[0101] • At t1, SII exceeds high_target_sii. The control system is now attempting to compensate by proposing changes.

[0102] • A suggested change has been made to the system, and the control system is waiting.

[0103] • At t2, SII lies between high_target_sii and dec_aim_target_sii. The control system again suggests new changes.

[0104] • This suggested change occurs, and the control system waits.

[0105] • At t3, SII is now lower than low_target_sii.

[0106] In the aforementioned scenario, the last change in the system caused the system to exceed the target at time t3. To correct this, in some examples, box 165 will involve reacting to and advising on the change in the same way as SII transitioning from the NONE state to low_target_sii. Therefore, in some such examples, box 165 will involve transitioning to the ref_gain_adj = DEC_REQ state.

[0107] The above example describes the case where SII changes from the NONE state to the INC_REQ state. A similar logic can be followed for cases where SII is below low_target_sii. In some such examples, box 165 will involve suggesting system changes. In some such examples, the control system will continue to suggest system changes in subsequent frames (with optional pauses between suggestions, as mentioned earlier) until SII is between inc_aim_target_sii and high_target_siiBetween. Before this condition is met, box 165 can correspond to the ref_gain_adj = DEC_REQ state.

[0108] Figure 5 Another graph showing the estimated SII over time is shown. Figure 5 Box 165 illustrates, based on some examples, how the control system can implement when and after the SII falls below the low target SII. Based on knowledge of how the system will react and how the SII will tend to fluctuate, Figure 5 shown inc_aim_target_sii Values ​​can be compared with target_sii They can be the same, or they can be different. For example, in some systems, it may be known that SII might be overestimated compared to the true SII due to the slow adaptation of noise estimation. In these cases, inc_aim_target_sii, slightly higher than target_sii (but lower than high_target_sii), might be preferred. Once SII falls between inc_aim_target_sii and high_target_sii (at t2), in some examples, the system state will return to the reference state. Figure 3B The described behavior. In some implementations, inc_aim_target_sii can be the same as target_sii.

[0109] An example of a system modification that the control system might suggest is increasing the volume. The volume increase can be specified such that the decrease is proportional to the difference between the current SII and the target SII. In one implementation, the control system might suggest a gain change in dB such that:

[0110] In the aforementioned equations, Gain_inc_delta represents the tuning parameter. In one embodiment, this tuning parameter is 0.8, but in alternative embodiments, it can be 0.7, 0.75, 0.85, 0.9, etc. In the aforementioned equations, Inc_speed represents another tuning parameter. In one embodiment, this tuning parameter is 1.0, but in alternative embodiments, it can be 0.9, 0.95, 1.05, 1.1, etc.

[0111] Whenever the control system suggests a change to the system, it may optionally wait a few frames to allow the system to stabilize before checking the current System Inversion Module (SII) and suggesting the new change. The amount of time the control system waits between suggestions can be proportional to the time between SII and the system inversion. The degree of proximity of inc_aim_target_sii is proportional. In one implementation, the waiting time can be determined as follows:

[0112] In the aforementioned equation, gain_adj_holdon_frames represents the tuning parameters. In one implementation, the tuning factor is 96 for a 20 ms block length. In other implementations, the tuning factor can be 90, 92, 94, 98, or 100 for a 20 ms block length.

[0113] In some cases, the control system can determine that the following has occurred: • At t0, SII is within high_target_sii and low_target_sii.

[0114] • At t1, SII is lower than low_target_sii. The control system is now attempting to compensate by proposing changes.

[0115] • A suggested change has been made to the system, and the control system is waiting.

[0116] • At t2, SII is between low_target_sii and inc_aim_target_sii. The control system again suggests new changes.

[0117] • This suggested change occurs, and the control system waits.

[0118] • At t3, SII is now higher than high_target_sii.

[0119] In this scenario, the last change in the system caused the system to exceed the target SII. To correct this, Box 165 could involve reacting to and advising on the change in the same way that SII breaks out of the NONE state from high_target_sii. In some examples, Box 165 would involve transitioning to the ref_gain_adj = INC_REQ state.

[0120] In the description of how box 165 can be implemented so far, we have not addressed how target_sii, high_target_sii, and low_target_sii can be determined. According to some examples, the control system can implement an initialization process during which these values ​​are determined. In some instances, the initialization process can be initiated by the user, while in others, it can be initiated when the device used to provide teleconferencing is powered on. In some examples, the other operations described above are not performed during the initialization step, and are only performed when initialization is complete.

[0121] According to some examples, the control system can begin the initialization process by determining the average SII over a certain time interval, which may correspond to the number of audio frames or blocks. (As used herein, the terms "audio frame" and "audio block" have the same meaning.) This block period can be considered a tuning parameter, referred to herein as "sii_max_init_counter". In one example, this tuning parameter may be 50 blocks, while in other examples, it may be 40 blocks, 45 blocks, 55 blocks, 60 blocks, etc. After determining the average SII over a certain time interval, in some implementations, the control system sets the average SII to target_sii. In some such implementations, the high target SII and the low target SII can be determined as follows:

[0122] In the aforementioned equation, target_sii_range represents the tuning parameter. In one example, this tuning parameter could be 0.15, while in other examples, it could be 0.1, 0.2, etc.

[0123] During the initialization phase, there may be audio blocks with weak energy from the loudspeaker. In these blocks, the SII calculation will not indicate the true SII of the system and therefore should not be considered during the SII averaging process. In some implementations, the control system can determine whether the energy is weak by checking if the following condition is met:

[0124] In the aforementioned equations, `mono_ref_level` represents the speaker energy of the current block (in dB), i.e., the input to the control system, and `ref_th_alpha` represents the tuning parameter. In one example, this tuning parameter could be 4.0, while in other examples, it could be 3.0, 3.5, 4.5, 5.0, etc. In the aforementioned equations, `ref_th` represents the tuning parameter. In one example, this tuning parameter could be -30, while in other examples, it could be -20, -25, -35, -40, etc.

[0125] When compensation is performed, there may be a goal to ensure that the system volume does not fall below the user's initial volume. To support this goal, during initialization, the control system can determine the system's current volume and then set the current volume to our minimum volume. This current volume can be stored, for example, as "usr_ref_gain". When a volume reduction is suggested in box 165, in some implementations, the control system can determine whether to implement the suggestion by ensuring that implementing the suggestion does not result in the volume falling below the minimum volume.

[0126] During the initialization phase, in some systems, the estimated SII may slope upwards and may be invalid for the first few blocks of the audio session. This phenomenon may be caused by instability in speech level or noise estimation during the first few blocks of the audio session. In some implementations, the control system can ensure that the SII averaging process does not consider the first few blocks of the audio session. In some such implementations, the control system may ignore the first 20 blocks (assuming a block length of 20 ms), the first 22 blocks, the first 24 blocks, the first 26 blocks, the first 28 blocks, the first 30 blocks, etc.

[0127] In some instances, during the implementation of box 165, the user may attempt to manually change the volume. If this occurs, the control system may optionally take into account the volume change attempted by the user. In some such examples, the control system may interpret the user-selected volume as a new minimum volume (new usr_ref_gain) and a new target_sii. According to some implementations, the control system may restart the initialization phase when the user interacts with the system's volume gain.

[0128] Figure 6 This is a flowchart outlining another example of a method that can be performed by an apparatus or system such as those disclosed herein. As with other methods described herein, the blocks of method 600 need not be performed in the indicated order. In some embodiments, one or more blocks of method 600 may be performed simultaneously. Furthermore, some embodiments of method 600 may include more or fewer blocks than those shown and / or described. The blocks of method 600 may be performed by one or more devices, which may be (or may include) a control system, such as… Figure 1B The control system 110 shown and described above.

[0129] In this example, method 600 is a method for compensating for ambient noise during a teleconference. According to this example, box 605 involves estimating the current speech spectrum corresponding to the speech of a remote teleconference participant by a control system. In some examples, box 605 may correspond to... Figure 1C Box 150, and can be executed according to the description of box 150 in this document.

[0130] According to this example, box 610 relates to the current noise spectrum estimated by the control system, corresponding to the ambient noise in the local environment where the local teleconference participants are located. According to some examples, box 610 may correspond to... Figure 1C Box 155, and can be executed according to the description of box 155 in this document.

[0131] In this example, box 615 relates to the calculation of the current speech intelligibility index (SII) by the control system based at least in part on the current speech spectrum and the current noise spectrum. In some examples, box 615 may correspond to Figure 1C Box 160, and can be executed according to the description of box 160 in this document.

[0132] According to this example, box 620 relates to a control system determining, at least in part, whether to adjust the local audio system used by local conference call participants based on the current SII. According to some examples, the determination may involve evaluating the current SII based on one or more target SII parameters. In some examples, the determination may involve determining whether the current SII is within the target SII range. In some examples, method 600 may involve adjusting at least a portion of the local audio system by the control system in response to determining that adjustment should be made. According to some examples, box 620 may correspond to... Figure 1C Box 165, and can be performed according to the description of box 165 herein. In some such examples, box 620 may refer to... Figures 3B to 5 One or more of the example implementations of the described box 165.

[0133] In this example, box 625 relates to the control system updating at least one of one or more target SII parameters in response to user input corresponding to a change in playback volume.

[0134] In some examples, method 600 may involve determining a confidence value corresponding to the current noise spectrum. The confidence value can indicate the probability that the current input audio frame primarily corresponds to ambient noise. The confidence value can be compared with the values ​​referenced herein. Figure 2A and Figure 2B The confidence metric 207 described corresponds to this. In some examples, method 600 may involve determining whether to update one or more noise statistics based at least in part on the confidence value.

[0135] Based on some examples, the confidence value can be a wideband confidence value. Determining whether to adjust the local audio system can be based at least in part on the wideband confidence value.

[0136] However, in some examples, method 600 may involve determining a band-based confidence value for each of a plurality of frequency bands. Determining whether to adjust the local audio system may involve determining, at least in part, whether to update one or more noise statistics for each frequency band based on the band-based confidence value.

[0137] In some examples, estimating the current noise spectrum may involve estimating the echo coupling gain corresponding to the local loudspeaker playback captured by the local microphone system. In some such examples, the echo coupling gain may be determined by... Figure 2B and Figure 3A The echo coupling gain is estimated using an echo coupling gain estimator 202. According to some examples, estimating the echo coupling gain may involve determining the maximum band loudspeaker reference power for each band of the current audio frame and the previous N-1 audio frames, where N is an integer corresponding to the number of audio frames in the delay line. In some examples, estimating the echo coupling gain may involve tracking the minimum power for each band of the input microphone signal. In some examples, estimating the echo coupling gain may involve estimating the echo coupling gain for each band of the input microphone signal based at least in part on the maximum band loudspeaker reference power and the minimum power. Some disclosed examples involve determining whether to update the estimated echo coupling gain based at least in part on whether the change in the currently estimated echo coupling gain value relative to the most recently estimated echo coupling gain value exceeds a threshold amount, based on whether the estimated echo coupling gain is updated within a threshold time interval, or based on a combination thereof. In some such examples, the threshold time interval can be measured within an audio frame or audio block.

[0138] In some examples, method 600 may involve adjusting at least a portion of a local audio system by a control system in response to determining that an adjustment should be made to maintain the current SII within a target SII range. According to some examples, maintaining the current SII within the target SII range may involve increasing the loudspeaker playback volume of one or more audio frames until the current SII is greater than a first target SII and less than a high target SII. In some examples, the first target SII may be greater than an intermediate target SII and less than a high target SII. According to some examples, maintaining the current SII within the target SII range involves decreasing the loudspeaker playback volume of one or more audio frames until the current SII is less than a second target SII and greater than a low target SII. In some such examples, the second target SII may be less than an intermediate target SII and greater than a low target SII.

[0139] Here are some example implementations of enumeration (EEE): EEE1. A method for compensating for ambient noise during a teleconference, the method comprising: estimating a current speech spectrum corresponding to the speech of a remote teleconference participant by a control system; estimating a current noise spectrum corresponding to ambient noise in a local environment in which a local teleconference participant is located by the control system; calculating a current speech intelligibility index (SII) by the control system based at least in part on the current speech spectrum and the current noise spectrum; determining a confidence value corresponding to the current noise spectrum, the confidence value indicating the likelihood that a current input audio frame primarily corresponds to ambient noise; and determining, by the control system and at least in part on the current SII and the confidence value, whether to adjust the local audio system used by the local teleconference participant, wherein the determination involves evaluating the current SII according to one or more target SII parameters.

[0140] EEE2. The method as described in EEE1, wherein the determination involves determining whether the current SII is within the range of the target SII.

[0141] EEE3. The method as described in EEE1 or EEE2 further includes, in response to determining that the adjustment should be made, adjusting at least a portion of the local audio system by the control system.

[0142] EEE4. The method as described in any one of EEE1 to 3 further includes updating at least one of the one or more target SII parameters in response to user input corresponding to a change in playback volume.

[0143] EEE5. The method as described in EEE1, wherein the confidence value is a wideband confidence value, and wherein determining whether to make the adjustment to the local audio system is based at least in part on the wideband confidence value.

[0144] EEE6. The method, as described in EEE4 or EEE5, further includes determining whether to update one or more noise statistics based at least in part on the confidence value.

[0145] EEE7. The method as described in EEE1 further includes determining a band-based confidence value for each of a plurality of frequency bands, and wherein determining whether to make the adjustment to the local audio system involves determining whether to update one or more noise statistics for each frequency band based at least in part on the band-based confidence value.

[0146] EEE8. The method of any one of EEE1 to 7, wherein estimating the current noise spectrum involves estimating the echo coupling gain corresponding to the local loudspeaker playback captured by the local microphone system.

[0147] EEE9. The method as described in EEE8, wherein estimating the echo coupling gain involves: determining the maximum band amplifier reference power for each band of the current audio frame and the previous N-1 audio frames, where N is an integer corresponding to the number of audio frames in the delay line; tracking the minimum power for each band of the input microphone signal; and estimating the echo coupling gain for each band of the input microphone signal based at least in part on the maximum band amplifier reference power and the minimum power.

[0148] EEE10. The method as described in EEE9 further includes determining whether to update the estimated echo coupling gain based at least in part on whether the currently estimated echo coupling gain value has changed by more than a threshold amount from the most recently estimated echo coupling gain value, whether the estimated echo coupling gain has been updated within a threshold time interval, or a combination thereof.

[0149] EEE11. The method of any one of EEE1 to 10 further includes, in response to determining that the adjustment should be made, the control system adjusting at least a portion of the local audio system to maintain the current SII within the target SII range.

[0150] EEE12. The method as described in EEE11, wherein maintaining the current SII within the target SII range involves increasing the loudspeaker playback volume of one or more audio frames until the current SII is greater than a first target SII and less than a high target SII.

[0151] EEE13. The method as described in EEE12, wherein the first target SII is greater than the intermediate target SII and less than the high target SII.

[0152] EEE14. The method as described in EEE12 or EEE13, wherein maintaining the current SII within the target SII range involves reducing the loudspeaker playback volume of one or more audio frames until the current SII is less than a second target SII and greater than a low target SII.

[0153] EEE15. The method as described in EEE14, wherein the second target SII is less than the intermediate target SII and greater than the low target SII.

[0154] EEE16. An apparatus configured to perform the method as described in any one of EEE1 to 15.

[0155] EEE17. A system configured to perform the method as described in any one of EEE1 to 15.

[0156] EEE18. One or more non-transitory media having instructions stored thereon for controlling one or more devices to perform a method as described in any one of EEE1 to 15.

[0157] Some aspects of this disclosure include a system or apparatus configured (e.g., programmed) to perform one or more examples of the disclosed methods, and a tangible computer-readable medium (e.g., a disk) storing code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems may be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including embodiments of the disclosed methods or steps thereof. Such a general-purpose processor may be or include a computer system including input devices, memory, and a processing subsystem programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data asserted thereto.

[0158] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed and otherwise configured) to perform necessary processing on (multiple) audio signals, including execution of one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed system (or elements thereof) may be implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor that may include input devices and memory), programmed with software or firmware and / or otherwise configured to perform any of a variety of operations, including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the system of the invention are implemented as general-purpose processors or DSPs configured (e.g., programmed) to execute one or more examples of the disclosed methods, and the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general-purpose processor configured to execute one or more examples of the disclosed methods may be coupled to input devices (e.g., a mouse and / or keyboard), memory, and a display device.

[0159] Another aspect of this disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) that stores code (e.g., an encoder executable to perform one or more examples of the disclosed methods or steps) for performing the disclosed methods or steps.

[0160] While specific embodiments and applications of this disclosure have been described herein, it will be apparent to those skilled in the art that many changes can be made to the embodiments and applications described herein without departing from the scope of this disclosure.

Claims

1. A method for compensating for ambient noise during a teleconference, the method comprising: The control system estimates the current speech spectrum corresponding to the speech of participants in the remote teleconference; The control system estimates the current noise spectrum corresponding to the ambient noise in the local environment where the local conference call participants are located; The current speech intelligibility index (SII) is calculated by the control system based at least in part on the current speech spectrum and the current noise spectrum. The control system determines, at least in part, whether to adjust the local audio system used by the local conference call participants, based on the current SII, wherein the determination involves evaluating the current SII according to one or more target SII parameters; as well as The control system updates at least one of the one or more target SII parameters in response to user input corresponding to changes in playback volume.

2. The method as described in claim 1, wherein, The determination involves determining whether the current SII is within the range of the target SII.

3. The method of claim 1 or claim 2, further comprising, in response to determining that the adjustment should be made, adjusting at least a portion of the local audio system by the control system.

4. The method of any one of claims 1 to 3, further comprising determining a confidence value corresponding to the current noise spectrum, the confidence value indicating the probability that the current input audio frame primarily corresponds to ambient noise.

5. The method of claim 4, wherein, The confidence value is a wideband confidence value, and the determination of whether to make the adjustment to the local audio system is based at least in part on the wideband confidence value.

6. The method of claim 4 or claim 5, further comprising determining whether to update one or more noise statistics based at least in part on the confidence value.

7. The method of claim 4, further comprising determining a band-based confidence value for each of a plurality of frequency bands, wherein, Determining whether to make the adjustment to the local audio system involves determining, at least in part, whether to update one or more noise statistics for each frequency band based on the frequency band-based confidence value.

8. The method according to any one of claims 1 to 7, wherein, Estimating the current noise spectrum involves estimating the echo coupling gain corresponding to the local loudspeaker playback captured by the local microphone system.

9. The method of claim 8, wherein, Estimating the echo coupling gain involves: Determine the maximum frequency band loudspeaker reference power for each frequency band of the current audio frame and the previous N-1 audio frames, where N is an integer corresponding to the number of audio frames in the delay line; The minimum power of each frequency band of the input microphone signal is tracked; and The echo coupling gain of each frequency band of the input microphone signal is estimated, at least in part, based on the maximum frequency band loudspeaker reference power and the minimum power.

10. The method of claim 9, further comprising determining whether to update the estimated echo coupling gain based at least in part on whether the currently estimated echo coupling gain value has changed by more than a threshold amount from the most recently estimated echo coupling gain value, whether the estimated echo coupling gain has been updated within a threshold time interval, or a combination thereof.

11. The method of any one of claims 1 to 10, further comprising, in response to determining that the adjustment should be made, adjusting at least a portion of the local audio system by the control system to maintain the current SII within the target SII range.

12. The method of claim 11, wherein, Maintaining the current SII within the target SII range involves increasing the loudspeaker playback volume of one or more audio frames until the current SII is greater than the first target SII and less than the high target SII.

13. The method of claim 12, wherein, The first target SII is greater than the intermediate target SII and less than the high target SII.

14. The method of claim 12 or claim 13, wherein, Maintaining the current SII within the target SII range involves reducing the loudspeaker playback volume of one or more audio frames until the current SII is less than the second target SII and greater than the low target SII.

15. The method of claim 14, wherein, The second target SII is smaller than the intermediate target SII and larger than the low target SII.

16. An apparatus configured to perform the method as claimed in any one of claims 1 to 15.

17. A system configured to perform the method as claimed in any one of claims 1 to 15.

18. One or more non-transitory media having instructions stored thereon for controlling one or more devices to perform the method as described in any one of claims 1 to 15.