Environmental noise compensation in teleconference

JP2026529473APending Publication Date: 2026-09-01DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026500701
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-17
Filing Date
2024-07-01
Publication Date
2026-09-01

Smart Images

  • Figure 2026529473000001_ABST
    Figure 2026529473000001_ABST
Patent Text Reader

Abstract

A method for compensating for ambient noise during a teleconference may include: the control system estimating the current voice spectrum corresponding to the voice of a remote teleconference participant; the control system estimating the current noise spectrum corresponding to ambient noise in the local environment where a local teleconference participant is located; the control system calculating a current speech clarity index (SII) based at least partially on the current voice spectrum and the current noise spectrum; the control system making a decision, at least partially based on the current SII, whether to adjust the local audio system used by the local teleconference participant, which includes evaluating the current SII according to one or more target SII parameters; and updating at least one of the one or more target SII parameters in response to user input corresponding to changes in playback volume.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 512,424, filed on 7 July 2023, and U.S. Provisional Patent Application No. 63 / 635,570, filed on 17 April 2024, all of which are invoked by reference in their entirety.

[0002] Technical field This disclosure relates to environmental noise compensation (ENC), particularly devices, systems, and methods for ENC in teleconferences. As used herein, the term “teleconference” encompasses both audio / video teleconferences and audio teleconferences. [Background technology]

[0003] Teleconference has become an important part of modern life. The ability to communicate clearly during teleconference is primarily based on speech intelligibility, which is partly based on the presence or absence of noise in the audio signal. Existing devices, systems, and methods for estimating speech intelligibility, noise estimation, and ENC are beneficial, but improved systems and methods are desired. [Overview of the Initiative]

[0004] At least some aspects of this disclosure may be implemented by method. Some such methods include compensating for ambient noise during a teleconference. For example, some methods may include a control system estimating a current speech spectrum corresponding to the speech of a remote teleconference participant, and a control system estimating a current noise spectrum corresponding to ambient noise in the local environment in which a local remote teleconference participant is located. Some methods may include a control system calculating a current speech intelligibility index (SII) based at least in part on the current speech spectrum and the current noise spectrum. Some methods may include a control system determining, at least in part on the current SII, whether to adjust the local audio system used by the local teleconference participant. Some methods may include a control system updating at least one of one or more target SII parameters in response to a user input corresponding to a change in playback volume.

[0005] The decision may involve evaluating the current SII according to one or more target SII parameters. In some examples, the decision may involve determining whether the current SII is within the target SII range. In some methods, the control system may involve adjusting at least a portion of the local audio system in response to a decision that adjustment should be performed.

[0006] Some methods may involve determining a confidence value corresponding to the current noise spectrum. The confidence value can, for example, indicate the likelihood that the current input audio frame primarily corresponds to ambient noise. In some examples, the confidence value may be a broadband confidence value. Deciding whether to adjust the local audio system may be based, at least partially, on the broadband confidence value. Some methods may involve determining a band-based confidence value for each of multiple frequency bands. In some examples, deciding whether to adjust the local audio system may involve deciding whether to update one or more noise statistics for each frequency band, at least partially based on the band-based confidence value. Some methods may involve deciding whether to update one or more noise statistics, at least partially based on the confidence value.

[0007] In some examples, estimating the current noise spectrum may involve estimating the echo coupling gain corresponding to local loudspeaker reproduction captured by the local microphone system. In some such examples, estimating the echo coupling gain may involve determining the maximum band loudspeaker reference power for each frequency band of the current audio frame and the previous N-1 audio frames, where N is an integer corresponding to the number of audio frames in the delay line. In such examples, estimating the echo coupling gain involves tracking the minimum power for each frequency band of the input microphone signal and estimating the echo coupling gain for each frequency band of the input microphone signal, at least in part, based on the maximum band loudspeaker reference power and the minimum power. Some methods may involve determining whether to update the estimated echo coupling gain based at least in part on (i) whether the current estimated echo coupling gain value has changed by more than a threshold amount from the most recently estimated echo coupling gain value, (ii) whether the estimated echo coupling gain has been updated within a threshold time interval, or (iii) both.

[0008] Some methods may involve adjusting at least a portion of the local audio system to maintain the current SII within the target SII range, depending on the decision that adjustments should be made. In some examples, maintaining the current SII within the target SII range may involve increasing the loudspeaker playback volume of one or more audio frames until the current SII is greater than the first target SII and less than the higher target SII. The first target SII may be greater than the median target SII and less than the higher target SII. In some examples, maintaining the current SII within the target SII range may involve decreasing the loudspeaker playback volume of one or more audio frames until the current SII is less than the second target SII and greater than the lower target SII. In some examples, the second target SII may be less than the median target SII and greater than the lower target SII.

[0009] Some or all of the operations, functions, and / or methods described herein may be executed by one or more devices in accordance with instructions (e.g., software) stored on one or more non-temporary media. Such non-temporary media may include, but are not limited to, memory devices such as random-access memory (RAM) devices and read-only memory (ROM) devices, as described herein. Accordingly, novel aspects of some of the subjects described herein may be implemented on non-temporary media storing software.

[0010] At least some aspects of this disclosure may be implemented by an apparatus. For example, one or more devices may be capable of performing at least partially the methods disclosed herein. In some implementations, the apparatus may include an interface system and a control system. The control system may include one or more general-purpose single or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or a combination thereof. In some examples, the apparatus may be one of the audio devices referenced above. However, in some implementations, the apparatus may be another type of device, such as a mobile device, laptop, or server.

[0011] Details of one or more implementations of the subject matter described in this specification are shown in the accompanying drawings and the following description. Other features, embodiments, and advantages will become further apparent from the description, drawings, and claims. Note that the relative dimensions in the following drawings may not be to scale. [Brief explanation of the drawing]

[0012] [Figure 1A] Figure 1A shows an example of a sound source that may be captured by a local microphone during a teleconference.

[0013] [Figure 1B] Figure 1B is a block diagram showing examples of components of an apparatus capable of carrying out various embodiments of the present disclosure.

[0014] [Figure 1C] Figure 1C is a flowchart illustrating an example of a method that may be performed by an apparatus or system such as those disclosed herein.

[0015] [Figure 2A] FIG. 2A illustrates an exemplary block of a novel noise estimator according to some examples.

[0016] [Figure 2B] FIG. 2B illustrates a block of the main noise estimator of FIG. 2A in accordance with some disclosed implementations.

[0017] [Figure 3A] FIG. 3A illustrates a block of the echo coupling gain estimator of FIG. 2B in accordance with some disclosed implementations.

[0018] [Figure 3B] FIG. 3B shows a time course plot of the estimated SII according to one example.

[0019] [Figure 4] FIG. 4 shows a time course plot of the estimated SII according to another example.

[0020] [Figure 5] FIG. 5 shows another time course plot of the estimated SII.

[0021] [Figure 6] FIG. 6 is a schematic flow diagram outlining another example method that may be performed by an apparatus or system as disclosed herein.

[0022] Like reference numerals and designations in the various drawings indicate like elements. DETAILED DESCRIPTION OF EMBODIMENTS

[0023] Figure 1A shows examples of sound sources that may be captured by the local microphone during a teleconference. In these examples, the local voice 102 of the local teleconference participant, ambient noise 104 (also referred to as “background noise” or “ambient noise”), and loudspeaker-generated sound 106 (also referred to as “echo”) from the loudspeaker 108 are captured by the microphone 112. During the teleconference, the voice of the remote teleconference participant (also referred to as “far-end voice”) is reproduced by the loudspeaker 108. The term “remote teleconference participant” refers to a teleconference participant located at a location other than that of the local teleconference participant.

[0024] Figure 1A also shows an example of a noise estimator 114, an embodiment thereof disclosed herein. The noise estimator 114 can be implemented, for example, by an instance of the control system 110 described with reference to Figure 1B.

[0025] The voices of remote teleconference participants and the local voices 102 of local teleconference participants are typically time-division multiplexed. Cases of "double talk," where the far-end voice and local voice 102 overlap in time, are uncommon. In most cases, double talk occurs only when teleconference participants wish to interrupt. To prevent recaptured far-end voice from being sent back to the far end, echo management is typically used in the signal chain, and the echo signal 106 is often barely suppressed.

[0026] Reliable noise estimation and noise reduction techniques can help meeting participants better understand what other teleconference participants are saying. Various improved noise estimation and noise reduction techniques are disclosed herein.

[0027] Figure 1B is a block diagram showing examples of components of a device capable of carrying out various embodiments of the present disclosure. As with other figures provided herein, the types and number of elements shown in Figure 1B are provided for illustrative purposes only. Other implementations may include more elements, fewer elements, different types of elements, or combinations thereof. In some examples, device 101 may be or include a device configured to perform at least some of the methods disclosed herein, such as a smart audio device, laptop computer, cellular phone, tablet device, or smart home hub. In some such implementations, device 101 may be or include a server configured to perform at least some of the methods disclosed herein.

[0028] In this example, the device 101 includes interface systems 105 and 110 control systems. In some implementations, the control system 110 may be configured to perform at least partially the methods disclosed herein. In some implementations, the control system 110 may be configured to compensate for ambient noise during teleconference.

[0029] In some examples, the control system 110 may be configured to estimate the current speech spectrum corresponding to the voice of a remote teleconference participant. According to some examples, the control system 110 may be configured to estimate the current noise spectrum corresponding to the ambient noise in the local environment where a local teleconference participant is located. In some examples, the control system 110 may be configured to calculate the current Speech Intelligibility Index (SII) based at least partially on the current speech spectrum and the current noise spectrum. In some examples, the control system 110 may be configured to decide, at least partially on the current SII, whether to adjust the local audio system used by the local teleconference participant. In some examples, the decision may include evaluating the current SII according to one or more target SII parameters. According to some examples, the control system 110 may be configured to update at least one of one or more target SII parameters in response to user input corresponding to a change in playback volume.

[0030] The interface system 105 may include one or more network interfaces and / or one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). Depending on the implementation, the interface system 105 may include one or more wireless interfaces. The interface system 105 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 105 may include one or more interfaces between the control system 110 and a memory system, such as the optional memory system 115 shown in Figure 1B. However, the control system 110 may include a memory system in some examples.

[0031] The control system 110 may include, for example, a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic, and / or discrete hardware components.

[0032] In some implementations, the control system 110 may reside in more than one device. For example, part of the control system 110 may reside in a device within a certain environment (such as a laptop computer, tablet computer, or smart audio device), while another part of the control system 110 may reside in a device outside that environment, such as a server. In other examples, part of the control system 110 may reside in a device within an environment, while another part of the control system 110 may reside in one or more other devices within that environment.

[0033] Some or all of the methods described herein may be executed by one or more devices in accordance with instructions (e.g., software) stored on one or more non-temporary media. Such non-temporary media may include, but are not limited to, random-access memory (RAM) devices, read-only memory (ROM) devices, and other memory devices as described herein. One or more non-temporary media may reside, for example, in the optional memory system 115 and / or control system 110 shown in Figure 1B. Thus, various novel aspects of the subject matter described herein can be implemented on one or more non-temporary media on which software is stored. The software may include, for example, instructions for controlling at least one device to process audio data. The software may be executable by one or more components of a control system, such as the control system 110 in Figure 1B.

[0034] In some examples, the device 101 may include an optional microphone system 120, as shown in Figure 1B. The optional microphone system 120 may include one or more microphones. In some implementations, one or more microphones may be part of or associated with another device, such as a loudspeaker or a smart audio device.

[0035] In some implementations, the device 101 may include an optional loudspeaker system 125, as shown in Figure 1B. The optional loudspeaker system 125 may include one or more loudspeakers. Loudspeakers are sometimes referred to as "speakers" in this document. In some examples, at least some of the loudspeakers of the optional loudspeaker system 125 may be arbitrarily positioned. For example, at least some of the speakers of the optional loudspeaker system 125 may be positioned in locations that do not correspond to any standard predetermined speaker layout, such as Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4, Dolby 9.1, Hamasaki 22.2, etc. In some such examples, at least some of the loudspeakers of the optional loudspeaker system 125 may be positioned in a location convenient to the space (e.g., where there is space to accommodate the loudspeakers), but not necessarily in any standard predetermined loudspeaker layout.

[0036] In some implementations, the device 101 may include an optional sensor system 130, as shown in Figure 1B. The optional sensor system 130 may include a touch sensor system, a gesture sensor system, one or more cameras, and the like.

[0037] In some implementations, the device 101 may include an optional display system 135, as shown in Figure 1B. The optional display system 135 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some examples, the optional display system 135 may include one or more organic light-emitting diode (OLED) displays. In some examples where the device 101 includes the display system 135, the sensor system 130 may include a touch sensor system and / or gesture sensor system adjacent to one or more displays of the display system 135. According to some such implementations, the control system 110 may be configured to control the display system 135 to present a graphical user interface (GUI), such as a GUI, related to implementing one of the methods disclosed herein.

[0038] As mentioned above, reliable noise estimation and noise reduction techniques can help meeting participants better understand what other teleconference participants are saying. Recent developments, including neural network-based approaches, are helping to solve the noise reduction problem. Neural network-based methods can often remove many types of noise, provided that the neural network training process noise provides the neural network with sufficient exposure to each specific type of noise to be removed. However, training a neural network to remove all possible types of noise may be difficult, or even impossible.

[0039] Therefore, noise compensation remains a crucial aspect of providing acceptable audio for teleconferences. The basic principle of noise compensation is to adjust the playback volume up or down based on ambient noise detected by the microphone. In some cases, noise compensation may be applied band by band, using different gains applied to different frequency bands.

[0040] In noise compensation scenarios, a robust and stable stationary noise estimator is crucial. The quality of the stationary noise estimator generally determines the overall system performance. Basic functions of a stationary noise estimator include: A stationary noise estimator should estimate only "stationary" noise to prevent abrupt fluctuations in volume in the case of accidental (dynamic) noises such as a dog barking, tapping on a table, or keyboard strokes. As used in this context, stationary noise refers not only to noise with "strict stationarity" where the statistical features do not change over time, but also to noise with "broad stationarity" where the first moment and autocovariance do not change over time, and the second moment is always finite. Stationary noise is a type of "stationary process," a stochastic process in which the unconditional joint probability distribution does not change when time is shifted. A steady-state noise estimator should estimate only the "true" ambient noise of the local environment, not the noise within the replayed echo signal. For example, referring to Figure 1A, the steady-state noise estimator should estimate only the ambient noise 104 and not the loudspeaker playback or "echo" sound 106. If the estimated noise is actually primarily caused by the echo (i.e., echo noise is dominant), the ambient noise compensation (ENC) system will form a positive feedback loop, which will cause some gain adjustment, resulting in a maximum or minimum volume setting.

[0041] Figure 1C is a flowchart illustrating an example of a method that may be performed by an apparatus or system such as that disclosed herein. The blocks of Method 140, as with other methods described herein, are not necessarily performed in the order shown. In some implementations, one or more blocks of Method 140 may be performed simultaneously. Furthermore, some implementations of Method 140 may include more or fewer blocks than those illustrated and / or described.

[0042] According to this example, Method 140 is a method for compensating for ambient noise during a teleconference. Blocks of Method 140 may be executed by one or more devices, which may be (or may include) a control system, such as the control system 110 shown in Figure 1B and described above. According to this example, blocks of Method 140 are repeated for each new block of audio data.

[0043] In this example, Method 140 includes calculating the current Speech Intelligibility Index (SII) and, if necessary, making incremental changes in the system, such as changes in playback volume, so that the system remains within the Speech Intelligibility Index target. The system may include a device configured to provide a teleconference. In some examples, the system may include a laptop having a loudspeaker and a microphone, configured to provide a teleconference and used by a local participant during the teleconference. The output of the loudspeaker during the teleconference is an example of what is referred to in this specification as “voice of a remote teleconference participant.” Referring again to Figure 1A, the audio captured by the local microphone would include ambient noise 104 and the voice of the local teleconference participant 102. To a noise detector, the voice 102 may be considered a type of interference.

[0044] In the example shown in Figure 1C, block 145 includes acquiring a frame of audio data from a local microphone. In this example, block 150 includes estimating the speech spectrum corresponding to the speech 102 in the current frame of audio data from the microphone. In this example, block 155 includes calculating the current noise spectrum corresponding to the ambient noise 104 for the current frame of audio data from the microphone. In some examples, blocks 150 and 155 may be performed simultaneously. In some examples, the noise spectrum is based at least in part on historical data. In some examples, block 155 will continue to update the noise spectrum as new data frames are received. In some examples, the current noise spectrum determined by block 155 is for the current frame of audio data from the microphone and is used to determine the SII and other parameters. In this example, block 160 includes calculating the current SII, which is the SII corresponding to the current frame of audio data from the microphone. According to this example, block 165 includes determining whether to make adjustments to the local audio system used by local teleconference participants, at least in part, based on the current SII.

[0045] In some cases, Block 150 may include estimating a speech spectrum as described in the "Methods for Calculation of the Speech Intelligibility Index" (hereinafter referred to as ASA / ANSI S3.5-1997), first published in 1969 and revised in 1997 by the American National Standards Institute and the Acoustical Society of America, which is incorporated hereby by reference. For example, Block 150 may include estimating a speech spectrum as described in the section "Methods for determining input variables for SII calculation" on pages 11-13 of ASA / ANSI S3.5-1997. However, in some alternative cases, Block 150 may include estimating a speech spectrum according to one or more other methods.

[0046] In some alternative examples, block 150 may involve estimating the speech spectrum according to a modified version of what is described in ASA / ANSI S3.5-1997. Some such examples involve using “insertion gain” which differs somewhat from what is described in ASA / ANSI S3.5-1997. In ASA / ANSI S3.5-1997, with respect to an amplification or attenuation device worn by a listener, at a particular frequency, insertion gain is the decibel difference between the pure tone sound pressure level (SPL) at the eardrum when the amplification / attenuation device is in place and the pure tone SPL at the eardrum when the amplification / attenuation device is removed. Insertion gain is used to calculate the equivalent speech spectral level (see section 5.1.3, beginning on page 12 of ASA / ANSI S3.5-1997) and also to calculate the equivalent noise spectral level (see section 5.1.4, page 13 of ASA / ANSI S3.5-1997). For example, according to ASA / ANSI S3.5-1997, when the speech spectrum is measured at the center of the listener's head, the equivalent speech spectrum for a particular frequency band is calculated as the measured speech spectrum level for that frequency band plus the insertion gain for that frequency band.

[0047] Some disclosed examples involve mapping the equivalent SPL of playback to this insertion gain. In some examples, the current hardware / software gain applied to the loudspeaker is converted to the equivalent SPL at the listener's position. According to some examples, this conversion is assisted by tuning parameters during the tuning process for a specific device, such as a particular laptop model. In some examples, this equivalent SPL is subtracted by the nominal SPL associated with the "normal" audio spectrum. The final subtracted value is used as the insertion gain, and the same value is used for all frequency bands.

[0048] In some examples, block 155 may include calculating the current noise spectrum corresponding to the ambient noise 104 in the current frame of audio data from the microphone, in accordance with the method described in ASA / ANSI S3.5-1997, for example, page 13. However, in some disclosed examples, block 155 may include calculating the current noise spectrum in accordance with an alternative method.

[0049] In some such methods, block 155 involves calculating the current noise spectrum corresponding to a new noise estimator (NE), an example of which is described in more detail with reference to Figures 2A-3A. Some implementations of the new noise estimator include at least a first, i.e., main noise estimator and a second, i.e., auxiliary noise estimator. The second noise estimator may be configured to respond to noise level changes that are relatively faster than those of the first or main noise estimator.

[0050] In some examples, block 160 may include determining the current SII as a value between 0 and 1 based on the current speech spectrum estimated by block 150 and the current noise spectrum estimated by block 155, where 1 means very clear and 0 means not clear at all. In some examples, block 160 may include determining the current SII based on the current speech spectrum, the current noise spectrum, and the measured or assumed auditory threshold level.

[0051] In some such examples, the SII calculated in block 160 may be the SII metric described in ASA / ANSI S3.5-1997. In some such examples, the SII metric may be calculated as described in the section “Methods for calculating Speech Intelligibility Index” on pages 9–11 of ASA / ANSI S3.5-1997, which is incorporated specifically hereby by reference.

[0052] According to some examples, the current SII may be determined as follows:

number

[0053] Figure 2A shows an exemplary block of a novel noise estimator in a partial example. In this example, the noise estimator 200 includes a first noise estimator 210, also referred to herein as the main noise estimator 210, a second noise estimator 220, also referred to herein as the auxiliary noise estimator 220, and a fast tracking flag module 230. According to this example, the main noise estimator 210, the auxiliary noise estimator 220, and the fast tracking flag module 230 are implemented by an instance of the control system 110 described in relation to Figure 1B. As with other figures provided herein, the types and number of elements shown in Figure 2A are provided merely as examples. Other implementations may include more elements, fewer elements, different types of elements, different arrangements of elements, or combinations thereof.

[0054] In the example shown in Figure 2A, both the main noise estimator 210 and the auxiliary noise estimator 220 are configured to receive loudspeaker audio data 222 and microphone audio data 224, respectively, as described with reference to block 145 in Figure 1C, for example. In some examples, the loudspeaker audio data 222 may be, or include, a reference signal corresponding to what is being played by the local loudspeaker.

[0055] In this example, the main noise estimator 210 is configured to determine and output a noise estimate 208 and a reliability metric 207, also referred to here as the “confidence value,” based at least partially on loudspeaker audio data 222 and microphone audio data 224. In this example, the noise estimate 208 and the reliability metric 207 correspond to the current noise spectrum, which corresponds to the current input audio frame of the microphone audio data 224. The reliability metric 207 indicates the likelihood that the current input audio frame corresponds primarily to ambient noise, such as the environmental noise 104 described with reference to Figure 1A. In some examples, the main noise estimator 210 may be configured to determine the noise estimate 208, the reliability metric 207, or both, based at least partially on a fast tracking flag 228 from the fast tracking flag module 230. Exemplary blocks and functions of the main noise estimator 210 are described in more detail below with reference to Figures 2B and 3A.

[0056] In this example, the auxiliary noise estimator 220 is configured to determine and output a noise estimate 226 based at least partially on loudspeaker audio data 222 and microphone audio data 224, where the noise estimate 226 corresponds to the current input audio frame of the microphone audio data 224. According to some disclosed examples, the auxiliary noise estimator 220 is configured to respond to changes in the noise level and generate a response noise estimate 226 relatively faster than the main noise estimator 210, thereby increasing the response speed (decreasing the response time) of the noise estimator 200.

[0057] In this example, the fast tracking flag module 230 is configured to determine and output the fast tracking flag 228 based at least in part on the noise estimates 226 and 208. In some examples, the main noise estimator 210 may be configured to converge to a new noise level relatively quickly if the main noise estimator 210 receives the fast tracking flag 228.

[0058] In some examples, if the noise estimate 226 from the auxiliary noise estimator 220 indicates that the broadband noise power is greater by a threshold amount for a determined time interval than the broadband noise power indicated by the noise estimate 208 from the main noise estimator 210, the fast tracking flag module 230 is configured to set the value of the fast tracking flag 228 to its maximum value, for example, 1. In other words, the fast tracking flag module 230 may be configured to set the value of the fast tracking flag 228 to its maximum value, for example, 1, if it is determined that the difference between (1) the noise power estimated by the auxiliary noise estimator 220 and (2) the noise power estimated by the main noise estimator 210 is greater than a difference threshold for at least a first time threshold. According to some examples, if it is determined that the difference between (1) and (2) is less than a difference threshold for at least a second time threshold, the fast tracking flag module 230 may be configured to set the value of the fast tracking flag 228 to its minimum value, for example, 0. The first time threshold may or may not be equal to the second time threshold. In some examples, the power threshold may be in the range of 1 dB to 5 dB, e.g., 1 dB, 2 dB, 3 dB, 4 dB, or 5 dB. According to some examples, the time constant may be in the range of 0.5 seconds to 2 seconds, e.g., 0.5 seconds, 0.6 seconds, 0.7 seconds, 0.8 seconds, 0.9 seconds, 1.0 seconds, 1.1 seconds, 1.2 seconds, 1.3 seconds, 1.4 seconds, 1.5 seconds, 1.6 seconds, 1.7 seconds, 1.8 seconds, 1.9 seconds, or 2 seconds.

[0059] In some examples, the auxiliary noise estimator 220 may be constructed using different tuning parameters than those of the main noise estimator 210. In some alternative examples, the auxiliary noise estimator 220 may be based on a deep neural network. Such a noise estimator may have a very fast response.

[0060] Figure 2B shows a block of the main noise estimator of Figure 2A in one of the disclosed implementations. In this example, the main noise estimator 210 includes a voice activity detector (VAD) 201, an echo coupling gain estimator 202, a confidence calculation block 203, and a noise and confidence statistics calculation block 204. According to this example, the VAD 201, the echo coupling gain estimator 202, the confidence calculation block 203, and the noise and confidence statistics calculation block 204 are implemented by an instance of the control system 110 described with reference to Figure 1B. As with other figures provided herein, the types and number of elements shown in Figure 2B are provided for illustrative purposes only. Other implementations may include more elements, fewer elements, different types of elements, different arrangements of elements, or combinations thereof.

[0061] In this example, the VAD 201 is configured to detect local voice activity based on microphone audio data 224 and send the voice activity signal 211 to the echo coupling gain estimator 202, the confidence calculation block 203, and the noise and confidence statistics calculation block 204. In some examples, the echo coupling gain estimator 202, the confidence calculation block 203, and the noise and confidence statistics calculation block 204 are active only if the voice activity signal 211 indicates that there is no local voice activity.

[0062] In this example, the echo coupling estimator 202 is configured to estimate the coupling gain from loudspeaker playback to microphone capture for each of several frequency bands, based on microphone audio data 224 and loudspeaker audio data 222, and to provide the corresponding coupling gain estimate 205 to the confidence calculation block 203. In other words, referring to the scenario shown in Figure 1A, the echo coupling gain estimator 202 is configured to estimate the contribution of the echo 106 to the sound detected by the microphone 112 and present in the microphone audio data 224, and to provide the corresponding coupling gain estimate 205. In the example shown in Figure 2B, the loudspeaker audio data 222 is or includes a reference signal corresponding to what is being played by the local loudspeaker. Figure 3A shows further details of the echo coupling gain estimator 202 in some examples.

[0063] In this example, confidence calculation block 203 is configured to calculate the likelihood that ambient noise is dominant in the current input audio frame of microphone audio data 224 for each frequency band. According to this example, confidence calculation block 203 is configured to output confidence data 206 to noise and confidence statistics calculation block 204. In this example, confidence data 206 includes at least a confidence value corresponding to the current noise spectrum, where the confidence value indicates the likelihood that the current input audio frame is primarily ambient noise. According to this example, confidence data 206 also includes a binary value (e.g., either 0 or 1) indicating whether the current input audio frame is likely to be primarily ambient noise and should be processed accordingly.

[0064] In this example, the noise and reliability statistics calculation block 204 is configured to determine and output a noise estimate 208 and a reliability metric 207 based at least partially on microphone audio data 224 and reliability data 206. The reliability metric 207 indicates the likelihood of the current input audio frame, primarily corresponding to ambient noise such as environmental noise 104 as described with reference to Figure 1A. In some examples, the noise and reliability statistics calculation block 204 may be configured to determine the noise estimate 208, the reliability metric 207, or both, based at least partially on a fast tracking flag 228 (not shown in Figure 2B) from the fast tracking flag module 230 in Figure 2A. A detailed example of how the noise and reliability statistics calculation block 204 may function is provided below.

[0065] Figure 3A shows a block of the echo-coupling gain estimator of Figure 2B in one of the disclosed implementations. In this example, the echo-coupling gain estimator 202 includes a delay line module 301, a minimum follower 302, a threshold detector 303, a maximum follower 304, an update control block 305, and a subtraction node 312. In this example, the delay line module 301, the minimum follower 302, the threshold detector 303, the maximum follower 304, the update control block 305, and the subtraction node 312 are implemented by an instance of the control system 110 described with reference to Figure 1B. In this example, the echo-coupling gain estimator 202 operates only when there is no local voice activity, as determined by the VAD 201 (see Figure 3A). As with other figures provided herein, the types and number of elements shown in Figure 3A are provided for illustrative purposes only. Other implementations may include more elements, fewer elements, different types of elements, different arrangements of elements, or combinations thereof.

[0066] In the example shown in Figure 3A, the delay line module 301 is configured to receive loudspeaker audio data 222 which is a delay line of length N and is either a loudspeaker reference signal corresponding to or containing a loudspeaker reference signal being played back by a local loudspeaker. In some examples, N is in the range of 8 to 16 frames, for example, 8 frames, 9 frames, 10 frames, 11 frames, 12 frames, 13 frames, 14 frames, 15 frames, 16 frames, etc. According to some examples, each frame may be in the range of 10 milliseconds (ms) to 30 ms, for example, 10 ms, 12 ms, 14 ms, 16 ms, 18 ms, 20 ms, 22 ms, 24 ms, 26 ms, 28 ms, 30 ms, etc. According to this example, the input loudspeaker reference signal is in the form of frequency-banded loudspeaker reference power. In this example, the delay line module 301 is configured to perform a maximum value acquisition (MAX) operation across the entire delay line. The purpose of passing the loudspeaker reference signal through a delay line is to compensate for the natural delay between the loudspeaker reference signal and the corresponding microphone capture of the sound reproduced by one or more local loudspeakers, taking into account sound wave propagation delay, electrical circuit delay, software buffering delay, and the formation of any reverberation that may be present in the local room. In this example, the delay line module 301 outputs maximum power data 310, which indicates the maximum loudspeaker reference signal power for each frequency band for the current audio frame and the previous N-1 audio frames.

[0067] In this example, the minimum follower 302 is configured to track the minimum power for each input frequency band for each frame of microphone audio data 224 and to output minimum power data 311 indicating the minimum input power for each frequency band for each frame of microphone audio data 224. In Figure 3A, "B" indicates the number of frequency bands. In some examples, the time window size of the minimum follower 302 may be dynamically adjusted. For example (referring again to Figure 2A), in some implementations, when a fast tracking flag 228 is input to the main noise estimator 210, the minimum follower 302 shortens its window size to help the main noise estimator 210 converge to a solution faster. In one such example, the minimum follower 302 can shorten its window size from 1.0 sec to 250 ms. Other examples may include different starting window sizes, different shortening window sizes, or both.

[0068] In the example shown in Figure 3A, the threshold detector 303 is a simple threshold detector configured to determine when the power of the current frame of input microphone audio data 224 exceeds the tracked minimum power level by at least a threshold amount, and to provide a corresponding threshold detector output 313. The tracked minimum power level depends heavily on the sensitivity of the particular microphone. Therefore, the tracked minimum power level range can vary greatly from microphone to microphone. According to some examples, the threshold amount is in the range of 3 dB to 10 dB, for example, 3 dB, 4 dB, 5 dB, 6 dB, 7 dB, 8 dB, 9 dB, or 10 dB. In some examples, the threshold detector output 313 of the threshold detector 303 always has a value of 1 when the power of the current frame of input microphone audio data 224 exceeds the tracked minimum by at least a threshold amount. In some such cases, the threshold detector output 313 always has a value of 0 when the power of the current frame of input microphone audio data 224 does not exceed the tracked minimum by at least a threshold amount.

[0069] In this example, the subtraction node 312 is configured to produce the subtraction node output 315 by subtracting (a) the input power of the current frame of the input microphone audio data 224 from (b) the maximum power of the loudspeaker reference signal, as determined by the maximum power data 310 output by the delay line module 301. The subtraction node output 315 represents the potential coupling gain estimate.

[0070] In this example, the maximum follower 304 is configured to determine and output a coupling gain estimate 205 based on the subtraction node output 315 and the update control signal 317 from the update control module 305.

[0071] In this example, the update control module 305 is configured to determine whether to allow or disallow the subtraction node output 315 to be output by the maximum follower 304 as the current coupling gain estimate 205. In some such examples, the update control signal 317 controls whether the subtraction node output 315 is received by the maximum follower 304 at all. In this example, even if the echo coupling gain estimator 202 operates only when there is no localized voice activity, localized transient noise may still be present in the microphone signal 224. Whenever there is any transient noise in the frequency band, the gain calculated by the subtraction node 312 is inaccurate and should be removed.

[0072] However, it can be difficult to determine whether the current input audio frame contains localized transient noise or simply originates from a strong echo, in other words, from a loudspeaker playback without prior knowledge of the coupling gain. This coupling gain is what the echo coupling gain estimator 202 needs to estimate here.

[0073] This problem can be overcome by using the fact that the coupling gain generally does not change abruptly and continuously for at least N frames, where N is the length of the delay line of the delay line module 301. Thus, according to some examples, whenever a new maximum gain exists, the control system 110 allows the maximum follower 304 to track the new maximum gain only if two conditions are met: 1. The new maximum gain is within a predetermined range of previously tracked maximum gains, e.g., 1 dB, 2 dB, 3 dB, 4 dB. In some implementations, this condition may be ignored for the first update if no initial measured value is available (see below). 2. There is no new maximum gain within the previous N frames.

[0074] The aforementioned problems can also be overcome by considering the fact that coupling gain is primarily a characteristic of the devices used to participate in teleconferences, such as telephones, laptops, and conference endpoints. While coupling gain can be affected by the acoustic characteristics of the room in which the device exists, it is primarily determined by the industrial design of the device itself. Once the product is manufactured, it is possible to obtain this coupling gain, and in some cases, this coupling gain may be stored as an initial value for future use.

[0075] Example of a confidence level calculation block function This section includes an example of how confidence calculation block 203 in Figure 2B may be implemented. Given the coupling gain and the current audio frame of microphone audio data 224, confidence calculation block 203 can calculate the likelihood of the current audio frame, which is primarily ambient noise. If the coupling gain (in dB) in frequency band b is g b Therefore, the maximum reference power of 310 is X b Assume (in dB units). Estimated echo power Eb ^ (in dB units) can be expressed as follows:

number

[0076] Power Y is ambient noise. b The likelihood of the current audio frame of microphone audio data 224 having the following can be expressed as follows:

number

[0077] In the formula mentioned above, σ is the threshold value, and may be 4 dB, 5 dB, 6 dB, 7 dB, 8 dB, etc.

[0078] According to some examples, the confidence calculation block 203 uses the instruction flag I b (corresponding to frequency band b) can be sent to the noise and confidence statistics calculation block 204. In some such cases, the confidence calculation block 203 receives the indicator flag I b This can be determined by implementing the following set of conditions:

number

[0079] In the above formula, n b This represents the estimated noise average for frequency band b. Second condition

number

[0080] Noise and confidence statistics calculation This section includes an example of how the noise and reliability statistics calculation block 204 of FIG. 2B may be implemented. In the example shown in FIG. 3A, the input to the noise and reliability statistics calculation block 204 includes minimum power data 311 from the minimum follower 302. The minimum power data 311 indicates the minimum value of input power for each frequency band of each frame of microphone audio data 224. According to this example, the input to the noise and reliability statistics calculation block 204 is microphone signal power Y b and an indication flag I b , both of which correspond to frequency band b.

[0081] In order to prevent transient noise from being incorporated into control and improve estimation accuracy, in some examples, the noise and reliability statistics calculation block 204 incorporates Y b into the accumulation when I=1 and the following conditions are satisfied: b :

Math

[0082] In the foregoing formula, Y b represents the tracked minimum value of Y b , and β represents a threshold such as 8 dB, 9 dB, 10 dB, 11 dB, 12 dB, or the like. In some examples, the reliability statistics are updated only when the noise accumulator is being updated, which includes L b as an input. According to some examples, the output of the noise and reliability statistics calculation block 204 includes n b , the estimated ambient noise power for each band, and p b , a reliability value corresponding to the current ambient noise power estimate.

[0083] Possible use cases of output from the noise and confidence statistics calculation blocks In some implementations, the reliability value p b bThis may be used, for example, to generate a broadband confidence flag P and control the behavior of the noise compensation control logic, as follows:

number

[0084] System adjustment decision and application After the control system estimates the current SII, in some implementations, the control system will perform a multi-step process to determine what the target SII is, whether to modify the system, and, if so, how to modify the system to achieve this target SII. In some examples, this process corresponds to block 165 in Figure 1C. Therefore, this section describes the various operations that may be performed according to the various implementations of block 165.

[0085] In some cases, block 165 can calculate and apply one or more changes to the system so that the SII calculated in block 160 in Figure 1C falls within the range of the target SII. Such changes may include, but are not limited to, the following: □ Changing the hardware / software gain of the speakers (e.g., volume control). An increase in gain will map to an increase in SII. The same logic applies to a decrease. □ Changes in equalization (EQ). □ Changes in the volume leveler effect of Dolby audio processing (DAP).

[0086] In some cases, such control is only implemented when the broadband confidence index is high (e.g., 1), in which case there is a high level of confidence that the estimated noise spectrum is indeed from background noise. If the confidence indicator is low (e.g., 0) and none of the above controls are in the user settings, in some cases the control system will cause the user settings to be restored.

[0087] The following are various examples of how to ensure that SII remains within the range of the target SII. Figure 3B shows an example plot of the estimated SII over time. The estimated SII may correspond, for example, to the output of block 160 in Figure 1C. Figure 3A shows how SII may fluctuate over a period of time. During this period, in this example, block 165 does not suggest any system changes because SII falls within the high and low targets, indicated as "HIGH TARGET sii" and "LOW TARGET sii" in Figure 3B. In this case, block 165 is in a state referred to here as "ref_gain_adj=NONE," indicating that no gain adjustment is made to the audio being played back on the local system while providing the teleconference to local teleconference participants.

[0088] Figure 4 shows a time-series plot of the estimated SII using another example. Figure 4 shows what happens when SII increases beyond high_target_sii. Between the start time and t1, block 165 does not suggest any volume change (ref_gain_adj=NONE). At t1, the SII value breaks out of its range and goes beyond high_target_sii. In this case, block 165 proposes a system change and continues to execute it in subsequent frames (with optional pauses between proposals, as described below) until SII is between dec_aim_target_sii and low_target_sii. Until that is achieved, block 165 is in ref_gain_adj=INC_REQ.

[0089] Based on knowledge of how the system reacts and how SII tends to fluctuate, the dec_aim_target_sii value may be the same as or different from target_sii. For example, it may be known that in some systems, SII may be underestimated compared to the true SII due to slow noise estimation adaptability. In these cases, a dec_aim_target_sii slightly lower than target_sii (but higher than low_target_sii) may be preferable. When SII falls between dec_aim_target_sii and low_target_sii (TIME=t2), in this example, the system state returns to the behavior described with reference to Figure 3B. In this implementation, dec_aim_target_sii is the same as target_sii.

[0090] One possible change to the system is to reduce the volume. The volume reduction can be specified to be proportional to the difference between the current SII and our target SII. In one implementation, the control system can determine possible gain changes in dB units, for example, the following:

number

[0091] In the above formula, gain_dec_delta represents a first tuning parameter, which is 0.6 in one example, and dec_speed represents a second tuning parameter, which is 1.0 in one example. Other examples may include different tuning parameters. In some alternative examples, the first tuning parameter may be 0.5, 0.55, 0.65, 0.7, etc., and the second tuning parameter may be 0.9, 0.95, etc.

[0092] In some implementations, whenever the control system determines a possible change to the system, it may be configured to optionally wait for a measurable time interval in the input audio frame to allow the system to stabilize before checking the current SII and proposing a new change. The proposed change may or may not be implemented. For example, if a function corresponding to one or more disclosed methods is switched off or disabled, for example according to user input, in some examples the proposed change will not be implemented. In some such implementations, the amount of time the control system is set to wait between checking the current SII and proposing a new change may be proportional to how close the SII is to dec_aim_target_sii. In some implementations, the control system may first calculate the following equation:

number

[0093] Based on the value of Diff_to_target, the control system can calculate the wait_factor, for example, as follows:

number

[0094] In some such examples, the number of input audio frames corresponding to the waiting interval is equal to the following:

number

[0095] In the above formula, gain_adj_holdon_frames represents the tuning parameter. In one implementation, the tuning factor is 96 for a block length of 20 ms. In other implementations, the tuning factor may be 90, 92, 94, 98, or 100 for a block length of 20 ms.

[0096] In some cases, the following may occur: □At t0, SII is within the range of high_target_sii and low_target_sii. At t1, SII exceeds high_target_sii. The control system attempts to compensate for this by proposing a change. □The proposed change occurs in the system, and the control system goes into standby mode. □At t2, SII is between high_target_sii and dec_aim_target_sii. The control system proposes a new change again. □When this suggested change occurs, the control system goes into standby mode. In t3, SII is now below low_target_sii.

[0097] In the aforementioned case, the last change in the system causes the system to overshoot the target by time t3. To correct this situation, in some examples, block 165 includes reacting in the same way and proposing a change, as if SII had broken out of the NONE state into low_target_sii. Thus, in some such examples, block 165 would involve a transition to the ref_gain_adj=DEC_REQ state.

[0098] The above example illustrates the case where SII transitions from the NONE state to the INC_REQ state. Similar logic can be followed when SII falls below low_target_sii. In some such cases, block 165 includes proposing a system change. In some such cases, the control system continues proposing system changes in subsequent frames until SII is between inc_aim_target_sii and high_target_sii (with optional pauses between proposals, as mentioned above). Until this condition is met, block 165 may correspond to the ref_gain_adj=DEC_REQ state.

[0099] Figure 5 shows another plot of the estimated SII time series. Figure 5 shows how the control system might perform block 165 when and after SII falls below the low target SII, according to some examples. Based on knowledge of how the system reacts and how SII tends to fluctuate, the inc_aim_target_sii value shown in Figure 5 may be the same as or different from target_sii. For example, it may be known that in some systems, SII can be overestimated compared to the true SII due to slow noise estimation adaptability. In these cases, an inc_aim_target_sii slightly higher than target_sii (but lower than high_target_sii) may be preferable. When SII falls between inc_aim_target_sii and high_target_sii (TIME=t2), in some examples, the system state returns to the behavior described with reference to Figure 3B. In some implementations, inc_aim_target_sii may be the same as target_sii.

[0100] One example of a system modification that the control system might propose is increasing the volume. The volume increase can be specified to be proportional to the difference between the current SII and the target SII. In one implementation, the control system may propose gain changes in dB units as follows:

number

[0101] In the above formula, gain_inc_delta represents a tuning parameter. In one implementation, this tuning parameter is 0.8, but in alternative implementations, this tuning parameter may be 0.7, 0.75, 0.85, 0.9, etc. In the above formula, Inc_speeed represents another tuning parameter. In one implementation, this tuning parameter is 1.0, but in alternative implementations, this tuning parameter may be 0.9, 0.95, 1.05, 1.1, etc.

[0102] Optionally, whenever the control system proposes a change to the system, it may check the current SII and propose a new change by waiting a few frames to allow the system to stabilize. The amount of time the control system waits between proposals may be proportional to how close the SII is to inc_aim_target_sii. In one implementation, the waiting time may be determined as follows:

number

[0103] In the above formula, gain_adj_holdon_frames represents the tuning parameter. In one implementation, the tuning coefficient is 96 for a block length of 20 ms. In other implementations, the tuning coefficient may be 90, 92, 94, 98, or 100 for a block length of 20 ms.

[0104] In some cases, the control system can confirm that the following occurs: At t0, SII is within the range of high_target_sii and low_target_sii. At t1, SII falls below low_target_sii. The control system attempts to compensate for this by proposing a change. □The proposed change occurs in the system, and the control system goes into standby mode. At t2, SII is between low_target_sii and inc_aim_target_sii. The control system proposes a new change again. □When this suggested change occurs, the control system goes into standby mode. In t3, SII is now above high_target_sii.

[0105] In this case, the last change in the system causes the system to overshoot the target SII. To correct this, block 165 may include proposing a change in the same way as if the SII had broken through high_target_sii from the NONE state. In some examples, block 165 would involve a transition to the ref_gain_adj=INC_REQ state.

[0106] The preceding explanation of how block 165 may be implemented does not mention how target_sii, high_target_sii, and low_target_sii may be determined. In some examples, the control system may perform an initialization process when these values ​​are determined. In some examples, the initialization process may be started by the user, while in others, the initialization process may be started when the device used to provide the teleconference is powered on. In some examples, the other actions described above will not be performed during the initialization step, and will only be performed if the initialization is complete.

[0107] In some examples, the control system can initiate the initialization process by determining the average SII over a time interval that can correspond to a number of audio frames or blocks. (As used here, the terms “audio frame” and “audio block” have the same meaning. The duration of a block may be considered as a tuning parameter referred to as “sii_max_init_counter”. In one example, this tuning parameter may be 50 blocks, but in other examples, this tuning parameter may be 40 blocks, 45 blocks, 55 blocks, 60 blocks, and so on.) After determining the average SII over a time interval, in some implementations, the control system will set this average SII to target_sii. In some such implementations, the high target SII and low target SII may be determined as follows:

number

[0108] In the formula above, target_sii_range represents the tuning parameter. In one example, this tuning parameter may be 0.15, while in other examples, it may be 0.1, 0.2, etc.

[0109] During the initialization phase, there may be audio blocks in the loudspeaker that have non-significant energy. For these blocks, the SII calculation does not represent the true SII of the system and therefore should not be considered during the averaged SII procedure. In some implementations, the control system may determine whether the energy is non-significant by checking whether the following conditions are met:

number

[0110] In the above formula, mono_ref_level represents the current block speaker energy in dB, the input to the control system, and ref_th_alpha represents the tuning parameter. In one example, this tuning parameter may be 4.0, and in other examples, it may be 3.0, 3.5, 4.5, 5.0, etc. In one example, ref_th represents the adjustment parameter. In one example, this tuning parameter may be -30, and in other examples, it may be -20, -25, -35, -40, etc.

[0111] When compensation is performed, one possible goal is to ensure that the system does not go below the user's initial volume level. To support this goal, during initialization, the control system may determine the system's current volume level and then set the current volume level as the minimum volume level. If a volume reduction is proposed in block 165, in some implementations, the control system may decide whether to implement the proposal by ensuring that implementing the proposal does not cause the volume to go below the minimum volume level.

[0112] During the initialization phase, in some systems, the estimated SII may ramp up, and may be invalid for the first few blocks of an audio session. This phenomenon can be caused by instability in the audio level or noise estimation during the first few blocks of an audio session. In some implementations, the control system may ensure that the SII averaging process does not consider the first few blocks of an audio session. In some such implementations, the control system may ignore the first 20 blocks (assuming a block length of 20 ms), the first 22 blocks, the first 24 blocks, the first 26 blocks, the first 28 blocks, the first 30 blocks, and so on.

[0113] In some examples, during the execution of block 165, the user may attempt to manually change the volume. If this occurs, the control system may optionally take into account the volume change attempted by the user. In some such examples, the control system may interpret the volume selected by the user as a new minimum volume (a new usr_ref_gain) and a new target_sii. According to some implementations, if the user has performed an operation on the system's volume gain, the control system may restart the initialization phase.

[0114] Figure 6 is a flowchart outlining another example of a method that may be performed by an apparatus or system such as that disclosed herein. The blocks of Method 600, as with other methods described herein, are not necessarily performed in the order shown. In some implementations, one or more blocks of Method 600 may be performed simultaneously. Furthermore, some implementations of Method 600 may include more or fewer blocks than those illustrated and / or described. The blocks of Method 600 may be performed by one or more devices that may be (or include) a control system such as the control system 110 shown in Figure 1B.

[0115] In this example, method 600 is a method for compensating for ambient noise during a teleconference. According to this example, block 605 includes the control system estimating the current voice spectrum corresponding to the voices of remote teleconference participants. In some examples, block 605 may correspond to block 150 in Figure 1C and may be performed according to the present description of block 150.

[0116] In this example, block 610 includes the control system estimating the current noise spectrum corresponding to the environmental noise in the local environment where the local teleconference participants are located. In some examples, block 610 may correspond to block 155 in Figure 1C and may be carried out according to the present description of block 155.

[0117] In this example, block 615 includes the control system calculating the current speech clarity index (SII) based at least in part on the current speech spectrum and the current noise spectrum. Block 615 may correspond to block 160 in Figure 1C and may be performed in accordance with the present description of block 160.

[0118] In this example, block 620 includes a control system making a decision, at least partially based on the current SII, whether to adjust the local audio system used by local teleconference participants. According to some examples, the decision includes evaluating the current SII according to one or more target SII parameters. In some examples, the decision may include determining whether the current SII is within the target SII range. In some examples, method 600 may include the control system adjusting at least a portion of the local audio system in response to a decision that adjustment should be made. According to some examples, block 620 may correspond to block 165 in Figure 1C and may be performed in accordance with the description of block 165 herein. In some such examples, block 620 may include one or more exemplary implementations of block 165 described with reference to Figure 3B-5.

[0119] In this example, block 625 includes the control system updating at least one of one or more target SII parameters in response to user input corresponding to a change in playback volume.

[0120] In some examples, Method 600 may include determining a confidence value corresponding to the current noise spectrum. The confidence value may represent the current input audio frame likelihood, primarily corresponding to ambient noise. The confidence value may correspond to the reliability metric 207, as described herein with reference to Figures 2A and 2B. In some examples, Method 600 may include determining whether to update one or more noise statistics, at least partially based on the confidence value.

[0121] In some examples, the confidence value may be a broadband confidence value. The decision of whether or not to adjust the local audio system may be based, at least in part, on the broadband confidence value.

[0122] However, in some examples, Method 600 may involve determining band-based confidence values ​​for each of multiple frequency bands. Deciding whether to adjust the local audio system may involve deciding whether to update one or more noise statistics for each frequency band, at least in part, based on the band-based confidence values.

[0123] In some such examples, the echo coupling gain may be estimated by the echo coupling gain estimator 202 in Figures 2B and 3A. In some examples, estimating the current noise spectrum may involve estimating the echo coupling gain corresponding to the local loudspeaker reproduction captured by the local microphone system. In some such examples, the echo coupling gain may be estimated by the echo coupling gain estimator 202 in Figures 2B and 3A. According to some examples, estimating the echo coupling gain may involve determining the maximum band loudspeaker reference power for each frequency band of the current audio frame and the previous N-1 audio frames, where N is an integer corresponding to the number of audio frames in the delay line. In some examples, estimating the echo coupling gain may involve tracking the minimum power for each frequency band of the input microphone signal. In some examples, estimating the echo coupling gain may involve estimating the echo coupling gain for each frequency band of the input microphone signal, at least in part, based on the maximum band loudspeaker reference power and the minimum power. Some of the disclosed examples include making a decision on whether to update the estimated echo coupling gain on at least part of: (i) whether the current estimated echo coupling gain value has changed by more than a threshold amount from the most recently estimated echo coupling gain value; (ii) whether the estimated echo coupling gain has been updated within a threshold time interval; or (iii) a combination thereof. In some of these examples, the threshold time interval may be measured in an audio frame or an audio block.

[0124] In some examples, method 600 may include, in response to a decision that adjustment should be made, the control system adjusting at least a portion of the local audio system to maintain the current SII within the target SII range. According to some examples, Maintaining the current SII within the target SII range may involve increasing the loudspeaker playback volume of one or more audio frames until the current SII is greater than the first target SII and less than the high target SII. In some examples, the first target SII may be greater than the median target SII and less than the high target SII. According to some examples, maintaining the current SII within the target SII range may involve decreasing the loudspeaker playback volume of one or more audio frames until the current SII is less than the second target SII and greater than the low target SII. In some such examples, the second target SII may be less than the median target SII and greater than the low target SII.

[0125] The following are some enumerated example embodiments (IEEE):

[0126] EEE1. How to compensate for ambient noise during teleconference: A control system estimates the current audio spectrum corresponding to the voices of remote teleconference participants; A step in which the control system estimates the current noise spectrum corresponding to the environmental noise in the local environment where the local teleconference participants are located; A step in which the control system calculates the current speech clarity index (SII) based at least partially on the current speech spectrum and the current noise spectrum; A step of determining a confidence value that corresponds to the current noise spectrum and represents the likelihood of the current input audio frame, which primarily corresponds to ambient noise; A decision on whether to adjust a local audio system used by local teleconference participants, comprising the steps of a control system making a decision, at least in part, based on the current SII and a confidence value, which includes evaluating the current SII according to one or more target SII parameters.

[0127] In the IEEE2.IEEE1 method, the determination includes determining whether the current SII is within the target SII range.

[0128] EEE3. In response to a decision that adjustments should be made in accordance with IEEE1 or IEEE2 methods, the control system further includes the step of adjusting at least a portion of the local audio system.

[0129] One of the methods of EEE4.IEEE1-3 further includes the step of updating at least one of one or more target SII parameters in response to a user input corresponding to a change in playback volume.

[0130] In the EEE5.IEEE1 method, the confidence value is the broadband confidence value, and the decision of whether or not to adjust the local audio system is based at least in part on the broadband confidence value.

[0131] EEE6. The method further includes determining whether to update one or more noise statistics based at least in part on confidence values, using an IEEE4 or IEEE5 method.

[0132] The EEE7.IEEE1 method further includes the step of determining a band-based confidence value for each of multiple frequency bands, and the step of deciding whether to adjust the local audio system includes the step of deciding whether to update one or more noise statistics for each frequency band, at least in part on the band-based confidence value.

[0133] In any one of the methods of EEE8.IEEE1-7, the step of estimating the current noise spectrum includes the step of estimating the echo coupling gain corresponding to local loudspeaker playback captured by the local microphone system.

[0134] In the IEEE9.IEEE8 method, the step for estimating the echo coupling gain is: A step of determining the maximum band loudspeaker reference power for each frequency band of the current audio frame and the previous N-1 audio frames, where N is an integer corresponding to the number of audio frames in the delay line; Steps include tracking the minimum power for each frequency band of the input microphone signal; and The process includes the step of estimating the echo coupling gain for each frequency band of the input microphone signal, at least in part, based on the maximum band loudspeaker reference power and minimum power.

[0135] In the IEEE10 and IEEE9 methods, the decision of whether or not to update the estimated echo coupling gain is made as follows: Whether the currently estimated echo coupling gain value has changed by more than a threshold amount from the most recently estimated echo coupling gain value. Whether the estimated echo coupling gain was updated within the threshold time interval, or This further includes steps that are based at least partially on those combinations.

[0136] In response to a decision that adjustment should be made in any one of the methods of EEE11.IEEE1-10, the control system further includes the step of adjusting at least a portion of the local audio system to maintain the current SII within the target SII range.

[0137] In the EEE12.IEEE11 method, the step of maintaining the current SII within the target SII range includes increasing the loudspeaker playback volume of one or more audio frames until the current SII is greater than the first target SII and less than the higher target SII.

[0138] In the EEE13.IEEE12 method, the first target SII is greater than the central target SII and less than the higher-level target SII.

[0139] EEE14. In the IEEE12 or IEEE13 method, the step of maintaining the current SII within the target SII range includes reducing the loudspeaker playback volume of one or more audio frames until the current SII is less than a second target SII and greater than a lower target SII.

[0140] In the EEE15.IEEE14 method, the second target SII is smaller than the central target SII and larger than the sub-target SII.

[0141] A device configured to perform one of the methods specified in EEE16.IEEE1-15.

[0142] A system configured to perform one of the methods specified in EEE17.IEEE1-15.

[0143] EEE18. One or more non-temporary media that store instructions for controlling one or more devices to perform any of the methods of IEEE1-15.

[0144] Some aspects of this disclosure include a system or device configured (e.g., programmed) to perform one or more examples of the disclosed method, and a tangible computer-readable medium (e.g., a disk) for storing code to perform one or more examples of the disclosed method or its steps. For example, some disclosed systems are programmable general-purpose processors, digital signal processors, or microprocessors programmed with software or firmware and / or configured to perform any of a variety of operations on data, which include or can include embodiments of the disclosed method or its steps. Such a general-purpose processor may be or include a computer system including an input device, memory, and a processing subsystem programmed (and / or otherwise configured) to perform one or more examples of the disclosed method (or its steps) in response to asserted data.

[0145] Some embodiments may be implemented as a constructable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed and otherwise configured) to perform necessary processing on an audio signal, including providing performance for one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed system (or its elements) may be implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include input devices and memory) programmed in software or firmware and / or otherwise configured to perform any of a variety of operations, including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the system of the present invention may be implemented as a general-purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed methods, and the system may also include other elements (e.g., one or more loudspeakers and / or one or more microphones). The general-purpose processor configured to perform one or more examples of the disclosed methods may be coupled with input devices (e.g., a mouse and / or keyboard), memory, and display devices.

[0146] Another aspect of this disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) that stores code (e.g., executable code for execution) for performing one or more examples of the disclosed method or its steps.

[0147] While specific embodiments and applications of the present disclosure are described herein, it will be apparent to those skilled in the art that many modifications are possible to the embodiments and applications described herein without departing from the scope of the present disclosure.

Claims

1. A method for compensating for ambient noise during teleconference: The control system estimates the current audio spectrum corresponding to the voices of remote teleconference participants; The control system estimates the current noise spectrum corresponding to the environmental noise in the local environment where the local teleconference participants are located; The control system calculates the current speech clarity index (SII) based at least partially on the current speech spectrum and the current noise spectrum; A step in which the control system makes a decision, at least in part, based on the current SII, whether to adjust the local audio system used by the local teleconference participant, the decision comprising evaluating the current SII according to one or more target SII parameters; and The control system updates at least one of the one or more target SII parameters in response to a user input corresponding to a change in playback volume; A method that includes this.

2. A method according to claim 1, wherein the determination comprises determining whether the current SII is within a target SII range.

3. A method according to claim 2, further comprising the step of the control system adjusting at least a portion of the local audio system in response to a decision that the adjustment should be made.

4. A method according to claim 1, further comprising the step of determining a confidence value corresponding to the current noise spectrum, wherein the confidence value indicates the likelihood of the current input audio frame, which primarily corresponds to ambient noise.

5. The method according to claim 4, wherein the confidence value is a broadband confidence value, and the step of deciding whether to adjust the local audio system is at least partially based on the broadband confidence value.

6. A method according to claim 4, further comprising the step of determining whether to update one or more noise statistics based at least in part on the confidence value.

7. A method according to claim 4, further comprising the step of determining a band-based confidence value for each of a plurality of frequency bands, wherein the step of determining whether to adjust the local audio system includes the step of determining whether to update one or more noise statistics for each frequency band, at least in part, based on the band-based confidence value.

8. A method according to claim 1, wherein the step of estimating the current noise spectrum includes the step of estimating the echo coupling gain corresponding to local loudspeaker playback captured by a local microphone system.

9. In the method according to claim 8, the step of estimating the echo coupling gain is: A step of determining the maximum band loudspeaker reference power for each frequency band of the current audio frame and the previous N-1 audio frames, where N is an integer corresponding to the number of audio frames in the delay line; Steps include tracking the minimum power for each frequency band of the input microphone signal, and A step of estimating the echo coupling gain for each frequency band of the input microphone signal, at least in part, based on the maximum band loudspeaker reference power and the minimum power; Methods that include...

10. In the method according to claim 9, the decision of whether to update the estimated echo coupling gain is made as follows: Whether the currently estimated echo coupling gain value has changed by more than a threshold amount from the most recently estimated echo coupling gain value. Whether the estimated echo coupling gain was updated within the threshold time interval, or A method that further includes steps performed based at least partially on those combinations.

11. A method according to claim 1, further comprising the step of the control system adjusting at least a portion of the local audio system in response to a decision that adjustment should be made to maintain the current SII within a target SII range.

12. A method according to claim 11, wherein the step of maintaining the current SII within a target SII range includes increasing the loudspeaker playback volume of one or more audio frames until the current SII is greater than a first target SII and less than a higher target SII.

13. The method according to claim 12, wherein the first target SII is greater than the central target SII and less than the higher-level target SII.

14. A method according to claim 12, wherein the step of maintaining the current SII within a target SII range includes reducing the loudspeaker playback volume of one or more audio frames until the current SII is less than a second target SII and greater than a lower target SII.

15. The method according to claim 14, wherein the second target SII is smaller than the central target SII and larger than the sub-target SII.

16. An apparatus configured to perform the method described in any one of claims 1 to 15.

17. A system configured to perform the method described in any one of claims 1 to 15.

18. One or more non-temporary storage media storing instructions for controlling one or more devices to perform the method described in any one of claims 1 to 15.