Method for reducing errors in an ambient noise compensation system

By measuring ambient noise at the listener's position and using critical distance analysis, the problem of noise source proximity ambiguity is solved, more accurate volume adjustment is achieved, and the noise compensation effect of audio equipment is improved.

CN114830681BActive Publication Date: 2025-08-08DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080085347.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-11-30
Filing Date
2020-12-08
Publication Date
2025-08-08
Estimated Expiration
2040-12-08

AI Technical Summary

Technical Problem

Existing audio equipment has a problem of noise source proximity ambiguity in noise compensation, resulting in noise estimation errors and the inability to accurately adjust the volume to adapt to the listener's actual noise environment.

Method used

By measuring the sound pressure level of the surrounding noise at the listener's position and using critical distance analysis and spectral attenuation time models, noise estimation errors are predicted, and the noise compensation method is adjusted to overcome proximity ambiguity and ensure the accuracy of volume adjustment.

Benefits of technology

It realizes more accurate noise compensation in different audio environments, avoids volume adjustment errors caused by noise source proximity problems, and provides a better audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114830681B_ABST
    Figure CN114830681B_ABST
Patent Text Reader

Abstract

Some noise compensation methods involve: receiving a microphone signal corresponding to ambient noise from a noise source location in or near an audio environment; determining or estimating a listener position in the audio environment; and estimating at least one critical distance, the critical distance being a distance from the noise source location at which the directly propagated sound pressure equals the diffuse-field sound pressure. Some examples involve estimating whether the listener position is within the at least one critical distance, and performing the noise compensation method for the ambient noise based at least in part on the estimation of whether the listener position is within the critical distance.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority from the following patent applications:

[0003] U.S. Provisional Patent Application No. 62 / 945,292, filed on December 9, 2019;

[0004] U.S. Provisional Patent Application No. 63 / 198,995, filed on November 30, 2020;

[0005] U.S. Provisional Patent Application No. 62 / 945,303, filed on December 9, 2019;

[0006] U.S. Provisional Patent Application No. 63 / 198,996, filed on November 30, 2020;

[0007] U.S. Provisional Patent Application No. 63 / 198,997, filed on November 30, 2020;

[0008] U.S. Provisional Patent Application No. 62 / 945,607, filed on December 9, 2019;

[0009] U.S. Provisional Patent Application No. 63 / 198,998, filed on November 30, 2020;

[0010] U.S. Provisional Patent Application No. 63 / 198,999, filed on November 30, 2020;

[0011] Each of the said U.S. Provisional Patent Applications is incorporated herein by reference in its entirety. Technical Field

[0012] The present disclosure relates to systems and methods for noise compensation. Background Art

[0013] Audio and video devices (including but not limited to televisions and associated audio equipment) are widely deployed.While existing systems and methods for controlling audio and video devices provide benefits, improved systems and methods would still be desirable.

[0014] Symbols and terminology

[0015] Throughout this disclosure, including in the claims, the terms "speaker," "loudspeaker," and "audio reproduction transducer" are used synonymously to refer to any sound-producing transducer (or set of transducers) driven by a single speaker feed. A typical set of headphones includes two speakers. A speaker may be implemented to include multiple transducers (e.g., a woofer and a tweeter), which may be driven by a single common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) may undergo different processing in different circuit branches coupled to different transducers.

[0016] Throughout this disclosure, including in the claims, expressions referring to “performing an operation on” a signal or data (e.g., filtering, scaling, transforming, or applying a gain to the signal or data) are used broadly to mean performing the operation directly on the signal or data or performing the operation on a processed version of the signal or data (e.g., a version of the signal that has undergone preliminary filtering or preprocessing before the operation is performed on it).

[0017] Throughout this disclosure, including in the claims, the expression "system" is used in a broad sense to refer to a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to a plurality of inputs, where the subsystem generates M of the inputs and the other XM inputs are received from external sources) may also be referred to as a decoder system.

[0018] Throughout this disclosure, including in the claims, the term "processor" is used in a broad sense to refer to a system or device that is programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, video, or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors that are programmed and / or otherwise configured to perform pipeline processing of audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.

[0019] Throughout this disclosure, including in the claims, the terms "couples" or "coupled" are used to mean either a direct or indirect connection. Thus, if a first device couples to a second device, that connection may be through a direct connection or through an indirect connection via other devices and connections.

[0020] As used herein, a "smart device" is an electronic device that can operate interactively and / or autonomously to some extent, and is typically configured to communicate with one or more other devices (or networks) via various wireless protocols such as Bluetooth, Zigbee, near-field communication, Wi-Fi, Light Fidelity (Li-Fi), 3G, 4G, 5G, etc. Several notable types of smart devices are smart phones, smart cars, smart thermostats, smart doorbells, smart locks, smart refrigerators, phablets and tablets, smart watches, smart bracelets, smart key chains, and smart audio devices. The term "smart device" may also refer to a device that exhibits some properties of ubiquitous computing, such as artificial intelligence.

[0021] In this document, the expression "smart audio device" is used to refer to a smart device, which is a single-purpose audio device or a multi-purpose audio device (for example, an audio device that implements at least some aspects of the virtual assistant functionality). A single-purpose audio device is a device (for example, a television (TV)) that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker and / or at least one camera) and is largely or primarily designed to implement a single purpose. For example, although a TV can generally play (and is considered to be able to play) audio from program material, in most cases, modern TVs run some kind of operating system on which applications (including applications for watching TV) run locally. In this sense, a single-purpose audio device with (multiple) speaker(s) and (multiple) microphone(s) is typically configured to run local applications and / or services to directly use (multiple) speaker(s) and (multiple) microphone(s). Some single-purpose audio devices can be configured to be combined together to enable the playback of audio on a certain zone or user-configured area.

[0022] A common type of multi-purpose audio device is an audio device that implements at least some aspects of a virtual assistant's functionality, although other aspects of the virtual assistant's functionality may be implemented by one or more other devices, such as one or more servers, with which the multi-purpose audio device is configured to communicate. Such a multi-purpose audio device may be referred to herein as a "virtual assistant." A virtual assistant is a device (e.g., a smart speaker or voice assistant integrated device) that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker and / or at least one camera). In some examples, a virtual assistant may provide the ability to use multiple devices (other than the virtual assistant) for applications that are cloud-enabled in some sense or that are otherwise not fully implemented in or on the virtual assistant itself. In other words, at least some aspects of the virtual assistant's functionality (e.g., speech recognition functionality) may be implemented (at least in part) by one or more servers or other devices, with which the virtual assistant may communicate via a network (e.g., the Internet). Virtual assistants may sometimes work together, for example, in a discrete and conditionally defined manner. For example, two or more virtual assistants may work together in the sense that one of them (e.g., the virtual assistant that is most confident of having heard the wake word) responds to the wake word. In some implementations, connected virtual assistants may form a constellation that may be managed by a host application, which may be (or implement) the virtual assistant.

[0023] In this document, "wake-up word" is used broadly to refer to any sound (e.g., a word spoken by a human or other sound) that a smart audio device is configured to wake up in response to detecting ("hearing") the sound (using at least one microphone included in or coupled to the smart audio device, or at least one other microphone). In this context, "waking up" means that the device enters a state of waiting (in other words, listening) for a voice command. In some instances, what may be referred to as a "wake-up word" herein may include more than one word, for example, a phrase.

[0024] In this document, the expression "wake-up word detector" refers to a device (or software including instructions for configuring the device to continuously search for alignment between real-time sound (e.g., speech) features and a trained model). Typically, a wake-up word event is triggered whenever the wake-up word detector determines that the probability of detecting the wake-up word exceeds a predefined threshold. For example, the threshold can be a predetermined threshold tuned to give a reasonable compromise between a false acceptance rate and a false rejection rate. After the wake-up word event, the device may enter a state (which may be referred to as an "awake" state or an "attention" state) in which the device listens for commands and passes received commands to a larger, more computationally intensive recognizer.

[0025] As used herein, the terms "program stream" and "content stream" refer to a collection of one or more audio signals, and in some instances, video signals, at least portions of which are intended to be heard together. Examples include a music selection, a movie soundtrack, a movie, a television program, the audio portion of a television program, a podcast, a live voice call, a synthesized voice response from an intelligent assistant, and the like. In some instances, a content stream may include multiple versions of at least a portion of an audio signal, for example, the same conversation in more than one language. In such instances, only one version of the audio data or portion thereof (e.g., a version corresponding to a single language) is intended to be reproduced at a time. Summary of the Invention

[0026] At least some aspects of the present disclosure may be implemented via one or more audio processing methods including, but not limited to, content stream processing methods. In some instances, method(s) may be implemented at least in part by a control system and / or via instructions (e.g., software) stored on one or more non-transitory media. Some such methods involve controlling a system by a first device and receiving a content stream comprising content audio data via a first interface system of a first device in an audio environment. In some examples, the first device may be a television or a television control module. Some such methods involve controlling a system by a first device and receiving a first microphone signal from a first device microphone system of the first device via a first interface system. Some such methods involve controlling a system by a first device and detecting ambient noise from a noise source location in or near an audio environment based at least in part on the first microphone signal.

[0027] Some such methods involve controlling a system by a first device to cause a first wireless signal to be transmitted from the first device via a first interface system to a second device in an audio environment. According to some embodiments, the first wireless signal may be transmitted via radio waves or microwaves. In some examples, the second device may be a remote control device, a smart phone, or a smart speaker. The first wireless signal may include instructions for causing the second device to record an audio segment, for example, via the second device microphone system. Some such methods involve controlling the system by the first device and receiving a second wireless signal from the second device via the first interface system. Some such methods involve determining, by the first device control system, a content stream audio segment time interval of the content stream audio segment. According to some embodiments, the second wireless signal may be transmitted via infrared waves.

[0028] Some such methods involve controlling, by the first device control system, and receiving, via the first interface system, a third wireless signal from the second device. The third wireless signal may include a recorded audio segment captured via the second device microphone. Some such methods involve determining, by the first device control system, a second device ambient noise signal at the second device location based at least in part on the recorded audio segment and the content stream audio segment. Some such methods involve implementing, by the first device control system, a noise compensation method on content audio data based at least in part on the second device ambient noise signal to produce noise-compensated audio data. In some examples, the method may involve controlling, by the first device control system, and providing, via the first interface system, the noise-compensated audio data to one or more audio reproduction transducers of the audio environment.

[0029] In some examples, the first wireless signal may include a second device audio recording start time or information for determining a second device audio recording start time. In some instances, the second wireless signal may indicate a second device audio recording start time. According to some examples, the method may involve controlling the system by the first device and receiving a fourth wireless signal from the second device via the first interface system. In some examples, the fourth wireless signal may indicate a second device audio recording end time. According to some examples, the method may involve determining a content stream audio segment end time based on the second device audio recording end time. In some instances, the first wireless signal may indicate a second device audio recording time interval.

[0030] According to some examples, the method may involve, for example, controlling, by the first device system, during a second device audio recording time interval and receiving, via the first interface system, a second microphone signal from the first device microphone system. In some examples, the method may involve detecting, by the first device system and based at least in part on the first microphone signal, a first device ambient noise signal corresponding to ambient noise from a noise source location. The noise compensation method may be based at least in part on the first device ambient noise signal. In some examples, the noise compensation method may be based at least in part on a comparison of the first device ambient noise signal with the second device ambient noise signal. According to some examples, the noise compensation method may be based at least in part on a ratio of the first device ambient noise signal to the second device ambient noise signal.

[0031] According to some examples, the method may involve controlling, by a first device control system, rendering the noise-compensated audio data to produce a rendered audio signal, and controlling, by the first device control system and providing, via a first interface system, the rendered audio signal to at least some of a set of audio reproduction transducers of an audio environment. In some implementations, at least one of the audio environment reproduction transducers may reside in the first device.

[0032] At least some alternative aspects of the present disclosure may be implemented via one or more audio processing methods, including but not limited to content stream processing methods. In some instances, the method(s) may be implemented at least in part by a control system and / or via instructions (e.g., software) stored on one or more non-transitory media. Some such methods involve receiving, by a control system and via an interface system, a microphone signal corresponding to ambient noise from a noise source location in or near an audio environment. Some such methods involve determining or estimating, by a control system, a listener position in an audio environment. Some such methods involve estimating, by a control system, at least one critical distance, the critical distance being the distance from the noise source location at which the directly propagated sound pressure equals the diffuse field sound pressure. Some such methods involve estimating whether a listener position is within at least one critical distance, and implementing a noise compensation method for the ambient noise based at least in part on at least one estimate of whether the listener position is within the at least one critical distance.

[0033] Some such methods may involve controlling an audio reproduction transducer system in an audio environment via a control system to reproduce one or more room calibration sounds, the audio reproduction transducer system comprising one or more audio reproduction transducers. In some examples, the one or more room calibration sounds may be embedded in content audio data received by the control system. Some such methods may involve receiving, by the control system and via an interface system, a microphone signal corresponding to a response of the audio environment to the one or more room calibration sounds, and determining, by the control system and based on the microphone signal, a reverberation time for each of a plurality of frequencies. Some such methods may involve determining or estimating an audio environment volume of the audio environment.

[0034] According to some examples, estimating at least one critical distance may involve calculating a plurality of estimated frequency-based critical distances based at least in part on the plurality of frequency-dependent reverberation times and the audio environment volume. In some examples, each of the plurality of estimated frequency-based critical distances may correspond to a frequency from the plurality of frequencies. In some examples, estimating whether the listener position is within the at least one critical distance may involve estimating whether the listener position is within each of the plurality of frequency-based critical distances. According to some examples, the method may involve transforming a microphone signal corresponding to ambient noise from the time domain to the frequency domain, and determining a frequency band ambient noise level estimate for each of a plurality of ambient noise frequency bands. According to some examples, the method may involve determining a frequency-based confidence level for each of the frequency band ambient noise level estimates. For example, each frequency-based confidence level may correspond to an estimate of whether the listener position is within each frequency-based critical distance. In some embodiments, each frequency-based confidence level may be inversely proportional to each frequency-based critical distance.

[0035] In some examples, implementing the noise compensation method may involve implementing a frequency-based noise compensation method for each ambient noise frequency band based on the frequency-based confidence level. In some instances, the frequency-based noise compensation method may involve applying a default noise compensation method for each ambient noise frequency band having a confidence level at or above a threshold confidence level. According to some embodiments, the frequency-based noise compensation method may involve modifying the default noise compensation method for each ambient noise frequency band having a confidence level below the threshold confidence level. For example, modifying the default noise compensation method may involve reducing a default noise compensation level adjustment.

[0036] According to some examples, the method may involve, by a control system and via an interface system, receiving a content stream comprising audio data. In some such examples, implementing a noise compensation method may involve applying the noise compensation method to the audio data to produce noise-compensated audio data. In some examples, the method may involve, by a control system and via an interface system, providing the noise-compensated audio data to one or more audio reproduction transducers of an audio environment.

[0037] In some examples, the method may involve rendering, by the control system, the noise-compensated audio data to produce a rendered audio signal, and providing, by the control system and via the interface system, the rendered audio signal to at least some of a set of audio reproduction transducers of the audio environment.

[0038] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, and the like. Thus, some innovative aspects of the subject matter described in this disclosure may be implemented via one or more non-transitory media having software stored thereon.

[0039] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of at least partially performing the methods disclosed herein. In some embodiments, the apparatus is or includes an audio processing system having an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or a combination thereof.

[0040] The details of one or more embodiments of the subject matter described in this specification are set forth in the following drawings and description. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. Note that the relative dimensions of the following figures may not be drawn to scale. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1A is a block diagram illustrating an example of a noise compensation system.

[0042] Figure 1B Another example of a noise compensation system is shown.

[0043] Figure 1C is a flow chart illustrating a method of scoring confidence in noise estimates using spectral decay time measurements according to some disclosed examples.

[0044] Figure 1D is a flow chart illustrating a method of using noise estimation confidence scores in a noise compensation process according to some disclosed examples.

[0045] Figure 2 is a block diagram illustrating an example of components of an apparatus capable of implementing various aspects of the present disclosure.

[0046] Figure 3 is a flow chart outlining one example of the disclosed method.

[0047] Figure 4A and Figure 4B Additional examples of noise compensation system components are shown.

[0048] Figure 4C It shows that Figure 4A and Figure 4B A timing diagram of an example of the operations performed by the noise compensation system is shown in FIG.

[0049] Figure 5 is a flow chart outlining one example of the disclosed method.

[0050] Figure 6 Additional examples of noise compensation systems are shown.

[0051] Figure 7A Is indicated by Figure 6 An example of an image of a signal received by a microphone is shown in FIG.

[0052] Figure 7B Shows the audio environment in different locations Figure 6 noise source.

[0053] Figure 8 An example of a floor plan of an audio environment, which in this example is a living space, is shown.

[0054] Throughout the drawings, like reference numbers and designations indicate like elements. DETAILED DESCRIPTION

[0055] The noise compensation system is configured to compensate for ambient noise within an audio environment, such as ambient noise. As used herein, the terms "ambient noise" and "environmental noise" refer to noise generated by one or more noise sources external to the audio playback system and / or the noise compensation system. In some examples, the audio environment may be a home audio environment, such as one or more rooms in a home. In other examples, the audio environment may be another type of environment, such as an office environment, a car environment, a train environment, a street or sidewalk environment, a park environment, etc.

[0056] Figure 1A An example of a noise compensation system is shown. In this example, the noise compensation system 100 is configured to adjust the level of an input audio signal 101 based on a noise estimate 108. According to this example, the noise compensation system 100 includes a loudspeaker 104, a microphone 105, a noise estimator 107, and a noise compensator 102. In some examples, the noise estimator 107 and the noise compensator 102 can be controlled via a control system (as described below with reference to FIG. Figure 2 The control system 210 described herein may be implemented, for example, according to instructions stored on one or more non-transitory storage media. As mentioned above, the terms "speaker," "loudspeaker," and "audio reproduction transducer" are used synonymously herein. As with the other figures provided herein, Figure 1A The types, quantities, and arrangements of the elements shown in FIG. 5 are provided as examples only. Other embodiments may include more, fewer, and / or different types, quantities, or arrangements of elements, such as more loudspeakers.

[0057] In this example, the noise compensator 102 is configured to receive an audio signal 101 from a file, a streaming service, etc. For example, the noise compensator 102 may be configured to apply a gain adjustment algorithm, such as a frequency-dependent gain adjustment algorithm or a wideband gain adjustment algorithm.

[0058] In this example, the noise compensator 102 is configured to send a noise-compensated output signal 103 to a loudspeaker 104. According to this example, the noise-compensated output signal 103 is also provided to a noise estimator 107 and is a reference signal for the noise estimator. In this example, a microphone signal 106 is also sent from the microphone 105 to the noise estimator 107.

[0059] According to this example, noise estimator 107 is a component configured to estimate the noise level in the environment including system 100. Noise estimator 107 can be configured to receive microphone signal 106 and calculate how much of microphone signal 106 is composed of noise and how much of microphone signal 106 is caused by playback from microphone 104. In some examples, noise estimator 107 can include an echo canceller. However, in some embodiments, noise estimator 107 can simply measure the noise when a signal corresponding to silence (a "quiet playback interval") is sent to microphone 104. In some such examples, a quiet playback interval can be an instance of an audio signal at or below a threshold level in one or more frequency bands. Alternatively or additionally, in some examples, a quiet playback interval can be an instance of an audio signal at or below a threshold level during a certain time interval.

[0060] In this example, noise estimator 107 is providing a noise estimate 108 to noise compensator 102. Depending on the particular implementation, noise estimate 108 may be a broadband estimate or a spectral estimate of the noise. In this example, noise compensator 102 is configured to adjust the level of the output of loudspeaker 104 based on noise estimate 108.

[0061] The loudspeakers of some devices, such as mobile devices, typically have quite limited capabilities. Therefore, the type of volume adjustment provided by system 100 will typically be limited by the dynamic range of such loudspeakers and / or speaker protection components (e.g., limiters and / or compressors). Noise compensation systems, such as noise compensation system 100, can apply gain that is either frequency-dependent or broadband.

[0062] While not yet commonplace in the consumer electronics market, the use of onboard microphones to measure and compensate for background noise has been demonstrated in home entertainment equipment. The primary reason this functionality hasn't been implemented is related to what this document will refer to as "noise source proximity ambiguity," the "proximity ambiguity problem," or simply the "proximity problem." In its simplest form, this problem arises from the fact that sound pressure level (SPL) is a measurement property that quantifies "how much sound is present" at a specific point in space. Because sound waves lose energy as they propagate through a medium, a measurement at one point in space is meaningless for all other points without a priori knowledge of the distances between these points and some properties of the transmission medium (in this case, air at room temperature). In anechoic space, it's straightforward to model these propagation losses using the inverse square law. This inverse square law doesn't apply to reverberant (real) rooms, so ideally, the reverberant properties of the physical space are also known in order to model propagation.

[0063] The proximity of a noise source to a listener is an important factor in determining the adverse effect of noise from that source on the listener's audibility and intelligibility of content. Measuring the sound pressure level via a microphone at an arbitrary location, such as on the casing of a television, is insufficient for determining the adverse effect of noise on a listener, as this microphone may perceive a very loud but distant noise source as the same sound pressure level as a quiet, nearby source.

[0064] The present disclosure provides various methods that can overcome at least some of these possible shortcomings, as well as devices and systems for implementing the presently disclosed methods. Some disclosed embodiments relate to measuring the SPL of ambient noise at the listener's position. Some disclosed embodiments relate to inferring the noise SPL at the listener's position from the level detected at any microphone position by knowing (or inferring) the proximity of the listener and the noise source to the microphone position. Various examples of the aforementioned embodiments are described below with reference to Figure 4 and below.

[0065] Some alternative embodiments involve predicting (e.g., on a per-frequency basis) how much error is likely to occur in the ambient noise estimate that does not involve a solution to the noise source proximity ambiguity problem. Figures 1B to 3 Some examples are described.

[0066] If a system does not implement one of the solutions described in the preceding paragraphs, some disclosed noise compensation methods may impose level adjustments on the output of the device that make the content reproduction too loud or too quiet for the listener.

[0067] Figure 1B Another example of a noise compensation system is shown. According to this example, the noise compensation system 110 includes a television 111, a microphone 112 configured to sample the acoustic environment (also referred to herein as the "audio environment") in which the noise compensation system 110 exists, and stereo loudspeakers 113 and 114. Figure 1B , but in this example, the noise compensation system 110 includes a noise estimator and a noise compensator, which may be the noise estimator and the noise compensator described above. Figure 1A In some examples, the noise estimator and the noise compensator may be controlled via a control system, such as a control system of the television 111 (which may be a control system of the television 111 described below). Figure 2 210 ), for example, according to instructions stored on one or more non-transitory storage media.

[0068] As with the other figures presented in this article, Figure 1B The types, quantities, and arrangements of the elements shown in FIG are provided as examples only. Other embodiments may include more, fewer, and / or different types and quantities of elements, for example, more loudspeakers. In some embodiments, reference Figures 1B to 1D The described noise compensation method can be implemented via the control system of a device other than a television, such as the control system of another device with a display (e.g., a laptop computer), the control system of a smart speaker, the control system of a smart hub, the control system of another device of an audio system, etc.

[0069] According to this example, a noise compensation system 110 is shown attempting to compensate for multiple noise sources to illustrate the aforementioned ambiguity regarding the proximity of noise sources to a listener. In this example, an audio environment 118 in which noise compensation system 110 is present also includes a listener 116 (assumed to be stationary in this example), a noise source 115 that is closer to a television 111 than listener 116, and a noise source 117 that is further away from television 111 than listener 116. In a highly damped room, without one of the disclosed methods for addressing or compensating for proximity, the noise compensation system may overcompensate for noise source 115. In a minimally damped room, without one of the disclosed methods for addressing or compensating for proximity, the noise compensation system may undercompensate for noise source 117 because noise source 117 is closer to listener 116 than to microphone 112.

[0070] In this example, the noise compensation system 110 is configured to implement a method based at least in part on a "critical distance" analysis. As used herein, the "critical distance" is the distance from a sound source at which the directly propagated sound pressure equals the diffuse field sound pressure. This property is frequency dependent and is typically given in terms of ISO standard octave or 1 / 3 octave band center frequencies. The critical distance is primarily a property of the volume (meaning the three-dimensional size, not the loudness) and reverberation of the audio environment (e.g., of a room), but is also affected by the directivity of the noise source. For a typical home living room with an omnidirectional source, the critical distance D is . c At 1kHz this is approximately 0.75m.

[0071] In a highly reverberant room, the noise compensation system 110 may provide adequate noise compensation despite failing to address the proximity problem. This is due to the fact that in a highly reverberant environment, the sound energy distribution throughout the room is nearly uniform beyond a critical distance.

[0072] In other words, in a highly reverberant room with a small critical distance, both listener 116 and television 111 will likely be outside the critical distance from the noise source. In this case, the reverberant sound dominates the direct sound, and the sound is relatively uniform regardless of the source distance and source position. Given this situation, it is unlikely that there will be a difference between the noise SPL measured at television microphone 112 and the noise SPL experienced by listener 116. This means that noise estimation errors due to proximity issues become less likely. Since both critical distance and reverberation time are frequency-dependent properties, this error probability also depends on frequency.

[0073] Unfortunately, most residential living rooms are not highly reverberant across all frequencies. In other words, at some frequencies, most residential living rooms may have a critical distance greater than 0.75 meters, and sometimes much greater than 0.75 meters. Therefore, it is possible that at some frequencies, the listener 116 and the television 111 may be within the critical distance. At such frequencies, a noise compensation system that has not yet addressed (or compensated for) the proximity issue will produce a noise estimate that is inaccurate for the noise level at the listener's position, and will therefore apply incorrect noise compensation.

[0074] Therefore, some disclosed embodiments involve predicting the probability of error due to proximity issues. To address this issue, existing functionality within some previously deployed devices can be leveraged to identify characteristics of the acoustic environment. At least some previously deployed devices that implement noise compensation will also have characteristics of a room acoustic compensation system. Using information already available from existing room acoustic compensation systems, frequency-dependent reverberation time (also known as spectral decay time) can be calculated. This is achieved by taking the system's impulse response (already calculated for the room acoustic compensation system) and dividing the impulse response into discrete frequency bands. The time from the peak of the impulse to the point where its magnitude decreases by 60 dB is the reverberation time for the frequency band.

[0075] After determining the spectral decay time, the spectral decay time and some knowledge of the room volume and source directivity can be used to infer a critical distance, and the control system can predict the probability of noise estimation error due to proximity issues based on the critical distance. If a small critical distance is predicted for a specific frequency bin (also referred to herein as a frequency range or band), then in some embodiments, this will produce a high confidence score (e.g., 1.0) for the ambient noise estimate in the frequency bin. According to some examples, the noise compensation system can then perform unconstrained noise compensation in the frequency bin. In some examples, unconstrained noise compensation can correspond to a "default" noise compensation that will be performed according to a noise compensation method in response to the ambient noise estimate in the frequency bin, for example, ensuring that the level of the playback audio exceeds the level of the ambient noise detected by microphone 112 by at least a threshold amount. In some examples, unconstrained noise compensation can correspond to a noise compensation method in which the output signal level of at least some frequency bands is not constrained according to the output signal level of other frequency bands and / or imposed thresholds.

[0076] In some embodiments, in frequency windows where the predicted critical distance is greater, this will result in lower confidence scores for those frequency windows. In some examples, the lower confidence scores result in the implementation of a modified noise compensation method. According to some such examples, the modified noise compensation method corresponding to the low confidence score can be a more conservative noise compensation method in which the level of the playback audio is raised less than that which would be raised according to the default method to reduce the likelihood of erroneously making large corrections.

[0077] According to some examples, a minimum (e.g., zero) confidence score can correspond to a minimum applied gain (e.g., a minimum difference between the reproduced audio level and the estimated ambient noise level), and a maximum (e.g., 1.0) confidence score can correspond to an unconstrained or "default" level adjustment for noise compensation. In some examples, confidence values between the minimum and maximum values can correspond to linear interpolations between the level adjustment corresponding to the minimum confidence score (e.g., minimum applied gain) and the "default" level adjustment for noise compensation.

[0078] In some embodiments, a minimum (e.g., zero) confidence score can correspond to a timbre-preserving noise compensation method, and a maximum (e.g., 1.0) confidence score can correspond to an unconstrained or "default" level adjustment for noise compensation. The term "timbre-preserving" can have various meanings as used herein. Broadly speaking, a "timbre-preserving" method is a method that at least partially preserves the frequency content or timbre of an input audio signal. Some timbre-preserving methods can completely or nearly completely preserve the frequency content of an input audio signal. The timbre-preserving method can involve constraining the output signal level of at least some frequency bands based on the output signal level of at least some other frequency bands and / or imposing thresholds. In some examples, the "timbre-preserving" method can involve constraining the output signal level of all non-isolated frequency bands, at least to some extent. (In some examples, if a frequency band is "isolated," only the audio in that frequency band contributes to the applied limiting gain.)

[0079] In some examples, the confidence value can be inversely proportional to the timbre preservation setting. For example, if the minimum confidence value is 0.0 and the maximum confidence value is 1.0, then a minimum (e.g., zero) confidence score can correspond to a timbre preservation setting of 100% or 1.0. In some examples, a timbre preservation setting of 0.50 can correspond to a confidence value of 0.5. In some such examples, a confidence value of 0.25 can correspond to a timbre preservation setting of 0.75.

[0080] For proximity issues to be considered insignificant in any given frequency window, the listener must be outside a critical distance for that frequency window. The critical distance for a particular frequency can be inferred from the reverberation time for that frequency using a statistical reverberation time model, for example, as follows:

[0081]

[0082] In Equation 1, D c represents the critical distance, Q represents the directivity factor of the noise source (in some embodiments, it is assumed to be omnidirectional), and V represents the volume of the room (e.g., in m 3 is the unit), and T represents the reverberation time measured in seconds, RT 60 RT 60 Defined as the time required for the amplitude of a theoretically perfect pulse to decay by 60dB.

[0083] In some examples, the volume of a room may be assumed to be a certain size, e.g., 60m, based on typical living room sizes. 3 . In some examples, the volume of a room can be determined based on input from the user at unboxing / setup time, e.g., via a graphical user interface (GUI). For example, the input can be numerical, based on the user's actual measurements or estimates. In some such implementations, the user can be presented with a set of "multiple-select" options via the GUI (e.g., "Is your room a large room, a medium-sized room, or a small room"). Each option can correspond to a different value of V.

[0084] In some embodiments, Equation 1 is solved for each of a plurality of frequency bins (e.g., for each frequency bin used by the noise compensation system 110). According to some examples, the confidence score can be generated by:

[0085] • Assume that the listener 116 will not be sitting within 2 meters of the TV 111.

[0086] If the predicted critical distance is equal to or less than 2 meters, the confidence score is set to 1.

[0087] As the critical distance increases, the confidence score decreases until it reaches the lower limit, where D c =5m and confidence level =0.

[0088] Alternative examples may involve alternative methods of determining a confidence score. For example, alternative methods may involve different assumptions about the proximity of the listener 116 to the television 111 and / or different critical distances for the lower limit, such as 4 meters, 4.5 meters, 5.5 meters, 6 meters, etc. Some embodiments may involve measuring or estimating the actual location of the listener 116 and / or the distance between the listener 116 and the television 111. Some embodiments may involve obtaining user input regarding the actual location of the listener 116 and / or the distance between the listener 116 and the television 111. Some examples may involve determining the location of a device, such as a cell phone or a remote control device, and assuming that the location of the device corresponds to the location of the listener.

[0089] According to various disclosed embodiments, the confidence scores described above represent the probability of error in noise compensation system 110's noise estimate. Given that in some embodiments there may be no way to distinguish between overestimation and underestimation, in some such embodiments, noise compensation system 110 may always assume that a noise estimation error is an overestimation. This assumption reduces the likelihood that noise compensation system 110 will erroneously apply too much gain to the audio reproduced by loudspeakers 113 and 114. Such embodiments are potentially advantageous because applying too much gain is typically a more perceptually noticeable failure mode than applying insufficient gain to sufficiently overcome ambient noise.

[0090] In some embodiments, if the confidence score is 1, the frequency-dependent gains calculated by the noise compensation system 110 are applied without constraint. According to some such embodiments, for all confidence values less than 1, these frequency-dependent gains are scaled down.

[0091] Figure 1C is a flow chart illustrating a method for scoring the confidence of a noise estimate using a spectral decay time measurement according to some disclosed examples. The figure illustrates the use of an impulse response, which, in some embodiments, may have been derived for the purpose of room acoustic compensation. According to the example, this impulse response is decomposed into discrete frequency bands corresponding to the frequency band in which the noise compensation system operates. The time it takes for each of these band-limited impulse responses to decay by 60 dB is the reverberation time RT60 for that frequency band.

[0092] Figure 1D is a flow chart illustrating a method for using noise estimation confidence scores in a noise compensation process according to some disclosed examples. For example, Figure 1C and Figure 1D The operations shown in FIG can be performed via a control system (as described below with reference to Figure 2As with other methods described herein, the blocks of method 120 and method 180 need not be performed in the order indicated. Additionally, such methods may include more or fewer blocks than shown and / or described. Figure 1C and Figure 1D In the example shown in , a box including a plurality of arrows indicates that the corresponding audio signal is divided into a number of frequency bins by the filter bank.

[0093] Figure 1C Method 120 may correspond to a "setup" mode, such as may occur when a television, audio device, or the like is first installed in an audio environment. In this example, block 125 involves causing one or more audio reproduction transducers to play a room calibration signal. Here, block 130 involves recording the room's impulse response to the room calibration signal via one or more microphones.

[0094] Here, block 135 involves transforming the impulse response from the time domain into the frequency domain: here, the corresponding audio signal is divided into a number of frequency bins by the filter bank. In this example, block 140 involves performing a decay time analysis and determining the reverberation time in seconds, RT 60 This analysis involves finding the peak of the impulse response for each band limit, counting the number of samples until the magnitude of the impulse response decays by 60 dB, and then dividing the number of samples by the sampling frequency in Hz. The result is the reverberation time RT60 in seconds for that band.

[0095] According to this example, block 145 involves determining a noise estimate confidence score for each of a plurality of frequency bins (e.g., for each frequency bin used by the noise compensation system 110). In some embodiments, block 145 involves solving Equation 1 for each of the frequency bins. Figure 1C Not shown, but in method 120 a value of V corresponding to the volume of the room is also determined, for example, based on user input, based on a room measurement or estimation process based on sensor input, or by using default values. According to some examples, a confidence score may be generated by assuming that listener 116 will not be sitting within 2 meters of television 111, with the confidence score being set to 1 when the predicted critical distance is equal to or less than 2 meters. As the critical distance increases, the confidence score may decrease, for example, to a lower limit where the critical distance is 5 meters and the confidence score is zero. Alternative examples may involve alternative methods of determining the confidence score. In some instances, the confidence score determined in block 145 may be stored in a memory.

[0096] In this example, Figure 1D The method 180 can be performed as follows: Figure 1CThe method of following corresponds to the "runtime" mode that occurs when a television, audio device, etc. is used on a daily basis. In this example, the echo cancellation block 155 involves receiving a microphone signal from the microphone 111 and also receiving an echo reference signal 150, which can be a speaker feed signal provided to an audio reproduction transducer of the audio environment. Here, block 160 involves generating a noise estimate for each of a plurality of frequency bins (also referred to herein as frequency bands) based on the output from the echo cancellation block 155.

[0097] In this example, the noise compensation scaling block 165 involves applying the confidence score determined in block 145 to provide appropriate scaling (if any) for the noise compensation gain to be applied based on the frequency-dependent noise estimate received from block 160. In some instances, the confidence score determined in block 145 may have been stored for later use, e.g., during runtime operation of method 180. For example, the scaling determined by the noise compensation scaling block 165 may be based on the above-described scaling with reference to FIG. Figure 1B to execute one of the examples described.

[0098] According to this example, block 170 involves determining a frequency-dependent gain based on the scaling value received from the noise compensation scaling block 165. Here, block 175 involves providing the noise-compensated output audio data to one or more audio transducers of the audio environment.

[0099] Figure 2 is a block diagram illustrating an example of components of an apparatus capable of implementing various aspects of the present disclosure. As with the other figures provided herein, Figure 2 , the type, quantity and arrangement of the elements shown in are provided as examples only. Other embodiments may include more, less and / or different types and quantities of elements. According to some examples, device 200 may be configured to perform at least some of the methods disclosed herein. In some embodiments, device 200 may be or may include one or more components of a television, an audio system, a mobile device (such as a cellular phone), a laptop computer, a tablet device, an intelligent speaker or another type of device. In some embodiments, device 200 may be or may include a television control module. Depending on the specific embodiment, the television control module may be integrated into the television or may not be integrated into the television. In some embodiments, the television control module may be a device separate from the television, and in some instances may be sold separately from the television or sold as an additional or optional device that the television purchased may include. In some embodiments, the television control module may be available from a content provider (such as a provider of television programs, movies, etc.).

[0100] According to some alternative embodiments, apparatus 200 may be or include a server. In some such examples, apparatus 200 may be or include an encoder. Thus, in some instances, apparatus 200 may be a device configured for use within an audio environment, such as a home audio environment, while in other instances, apparatus 200 may be a device configured for use in the "cloud," e.g., a server.

[0101] In this example, the device 200 includes an interface system 205 and a control system 210. In some embodiments, the interface system 205 can be configured to communicate with one or more other devices in the audio environment. In some examples, the audio environment can be a home audio environment. In other examples, the audio environment can be another type of environment, such as an office environment, a car environment, a train environment, a street or sidewalk environment, a park environment, etc. According to some embodiments, the size and / or reverberation of the audio environment can be assumed based on the audio environment type. For example, a default office size can be used for an office audio environment. For example, the audio environment type can be determined based on user input or based on the audio characteristics of the environment. In some embodiments, the interface system 205 can be configured to exchange control information and associated data with audio devices in the audio environment. In some examples, the control information and associated data can be related to one or more software applications being executed by the device 200.

[0102] In some embodiments, the interface system 205 can be configured to receive a content stream or to provide a content stream. The content stream may include audio data. The audio data may include but may not be limited to an audio signal. In some instances, the audio data may include spatial data such as channel data and / or spatial metadata. According to some embodiments, the content stream may include metadata about the dynamic range of the audio data and / or metadata about one or more noise compensation methods. For example, metadata about the dynamic range of the audio data and / or metadata about one or more noise compensation methods may have been provided by one or more devices (such as one or more servers) configured to implement a cloud-based service. For example, metadata about the dynamic range of the audio data and / or metadata about one or more noise compensation methods may have been provided by a device that may be referred to as an "encoder" in this document. In some such examples, the content stream may include video data and audio data corresponding to the video data. Some examples of encoder and decoder operations are described below.

[0103] The interface system 205 may include one or more network interfaces and / or one or more external device interfaces (e.g., one or more universal serial bus (USB) interfaces). According to some embodiments, the interface system 205 may include one or more wireless interfaces. The interface system 205 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 205 may include a control system 210 and a memory system (e.g., Figure 2 205). However, in some instances, the control system 210 may include a memory system. In some implementations, the interface system 205 may be configured to receive input from one or more microphones in the environment.

[0104] For example, the control system 210 may include a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.

[0105] In some embodiments, the control system 210 may be present in more than one device. For example, in some embodiments, a portion of the control system 210 may be present in a device within one of the environments described herein, and another portion of the control system 210 may be present in a device outside the environment such as a server, a mobile device (e.g., a smart phone or tablet computer). In other examples, a portion of the control system 210 may be present in a device within one of the environments described herein, and another portion of the control system 210 may be present in one or more other devices of the environment. For example, the control system functionality may be distributed across multiple smart audio devices of the environment, or may be shared by an orchestration device (e.g., a device that may be referred to as a smart home hub herein) and one or more other devices of the environment. In other examples, a portion of the control system 210 may be present in a device (e.g., a server) that implements a cloud-based service, and another portion of the control system 210 may be present in another device (e.g., another server, a storage device, etc.) that implements a cloud-based service. In some examples, the interface system 205 may also be present in more than one device.

[0106] In some implementations, the control system 210 can be configured to at least partially perform the methods disclosed herein.According to some examples, the control system 210 can be configured to implement a method of content streaming.

[0107] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. For example, one or more non-transitory media may reside in Figure 2 2 and / or in the control system 210. Thus, various innovative aspects of the subject matter described in this disclosure may be implemented in one or more non-transitory media having software stored thereon. For example, the software may include instructions for controlling at least one device to process a content stream, encode a content stream, decode a content stream, etc. For example, the software may be a program that can be used by a control system (e.g., Figure 2 Executed by one or more components of the control system 210).

[0108] In some examples, apparatus 200 may include Figure 2 Optional microphone system 220 is shown in FIG. Optional microphone system 220 may include one or more microphones. In some embodiments, one or more microphones may be part of or associated with another device (e.g., a speaker of a speaker system, a smart audio device, etc.). In some examples, device 200 may not include microphone system 220. However, in some such embodiments, device 200 may still be configured to receive microphone data from one or more microphones in the audio environment via interface system 210. In some such embodiments, a cloud-based implementation of device 200 may be configured to receive microphone data or noise metrics corresponding at least in part to the microphone data from one or more microphones in the audio environment via interface system 210.

[0109] According to some embodiments, the apparatus 200 may include Figure 2 Optional loudspeaker system 225 is shown in FIG. Optional loudspeaker system 225 may include one or more loudspeakers, which may also be referred to herein as "speakers," or more generally as "audio reproduction transducers." In some examples (e.g., cloud-based implementations), device 200 may not include loudspeaker system 225.

[0110] In some embodiments, the apparatus 200 may include Figure 2Optional sensor system 230 is shown in . Optional sensor system 230 may include one or more touch sensors, gesture sensors, motion detectors, etc. According to some embodiments, optional sensor system 230 may include one or more cameras. In some embodiments, the camera may be a stand-alone camera. In some examples, one or more cameras of optional sensor system 230 may be present in a smart audio device, which may be a single-purpose audio device or a virtual assistant. In some such examples, one or more cameras of optional sensor system 230 may be present in a television, a mobile phone, or a smart speaker. In some examples, device 200 may not include sensor system 230. However, in some such embodiments, device 200 may still be configured to receive sensor data from one or more sensors in the audio environment via interface system 210.

[0111] In some embodiments, the apparatus 200 may include Figure 2 Optional display system 235 is shown in . Optional display system 235 may include one or more displays, such as one or more light emitting diode (LED) displays. In some instances, optional display system 235 may include one or more organic light emitting diode (OLED) displays. In some examples, optional display system 235 may include one or more displays of a television. In other examples, optional display system 235 may include a laptop display, a mobile device display, or another type of display. In some examples in which device 200 includes display system 235, sensor system 230 may include a touch sensor system and / or a gesture sensor system proximate to one or more displays of display system 235. According to some such embodiments, control system 210 may be configured to control display system 235 to present one or more graphical user interfaces (GUIs).

[0112] According to some such examples, apparatus 200 may be or include a smart audio device. In some such implementations, apparatus 200 may be or include a wake-up word detector. For example, apparatus 200 may be or include a virtual assistant.

[0113] Figure 3 300 is a flowchart outlining an example of the disclosed method. As with other methods described herein, the blocks of method 300 do not have to be executed in the order indicated. In addition, such a method may include more or fewer blocks than shown and / or described.

[0114] Method 300 may be performed as follows: Figure 22 and described above. In some examples, the blocks of method 300 can be performed by one or more devices within the audio environment, for example, an audio system controller or another component of the audio system, such as a smart speaker, a television, a television control module, a smart speaker, a mobile device, etc. In some embodiments, the audio environment can include one or more rooms of a home environment. In other examples, the audio environment can be another type of environment, such as an office environment, a car environment, a train environment, a street or sidewalk environment, a park environment, etc. However, in alternative embodiments, at least some blocks of method 300 can be performed by a device (such as a server) that implements a cloud-based service.

[0115] In this embodiment, block 305 involves receiving, by the control system and via the interface system, a microphone signal corresponding to ambient noise from a noise source location in or near the audio environment. In some embodiments, the control system and the interface system may be Figure 2 The control system 210 and interface system 205 are shown in FIG and described above.

[0116] In this example, block 310 involves determining or estimating, by the control system, a listener position in the audio environment. According to some examples, block 310 may involve determining the listener position based on a default value of an assumed listener position, e.g., the listener is 2 meters in front of the television or other device or at least 2 meters in front of the television or other device, the listener is sitting on a piece of furniture with a known position referenced to the television or other device, etc. However, in some implementations, block 310 may involve determining the listener position based on user input, based on (e.g., from Figure 2 The listener position may be determined using sensor input from a camera of the sensor system 230 shown in FIG. 2 . Some examples may involve determining the position of a device such as a cell phone or remote control device and assuming that the device's position corresponds to the listener position.

[0117] According to this example, block 315 involves estimating, by the control system, at least one critical distance. As described elsewhere herein, the critical distance is the distance from the noise source location at which the directly propagated sound pressure equals the diffuse field sound pressure. In some examples, block 315 may involve retrieving the at least one estimated critical distance from a memory, Figure 1CThe results of the method or similar method are stored in the memory. Some such methods may involve controlling an audio reproduction transducer system in an audio environment via a control system to reproduce one or more room calibration sounds. The audio reproduction transducer system includes one or more audio reproduction transducers. Some such methods may involve receiving, by the control system and via an interface system, a microphone signal corresponding to a response of the audio environment to the one or more room calibration sounds. Some such methods may involve determining, by the control system and based on the microphone signal, a reverberation time for each of a plurality of frequencies. Some such methods may involve determining or estimating an audio environment volume of the audio environment (in other words, determining the size of the audio environment in cubic feet, cubic meters, etc.), for example, as disclosed elsewhere herein. According to some such examples, estimating at least one critical distance may involve calculating a plurality of estimated frequency-based critical distances based at least in part on the plurality of frequency-dependent reverberation times and the audio environment volume. Each of the plurality of estimated frequency-based critical distances may correspond to a frequency from the plurality of frequencies.

[0118] In this example, block 320 involves estimating whether the listener position is within at least one critical distance. According to some examples, block 320 may involve estimating whether the listener position is within each of a plurality of frequency-based critical distances. In some examples, method 300 may involve transforming a microphone signal corresponding to ambient noise from the time domain into the frequency domain and determining a frequency band ambient noise level estimate for each of a plurality of ambient noise frequency bands. According to some such examples, method 300 may involve determining a frequency-based confidence level for each of the frequency band ambient noise level estimates. For example, each frequency-based confidence level may correspond to an estimate or probability of whether the listener position is within each frequency-based critical distance. In some examples, each frequency-based confidence level may be inversely proportional to each frequency-based critical distance.

[0119] According to this embodiment, block 325 involves implementing a noise compensation method for the ambient noise based at least in part on at least one estimate of whether the listener position is within at least one critical distance. In some examples, block 325 may involve implementing a frequency-based noise compensation method for each ambient noise frequency band based on a frequency-based confidence level. According to some such examples, the frequency-based noise compensation method may involve applying a default noise compensation method for each ambient noise frequency band whose confidence level is at or above a threshold confidence level. In some instances, the threshold confidence level may be a maximum confidence level, e.g., 1.0. However, in other examples where the maximum confidence level is 1.0, the threshold confidence level may be another confidence level, e.g., 0.80, 0.85, 0.90, 0.95, etc.

[0120] In some examples, the frequency-based noise compensation method may involve modifying a default noise compensation method for each ambient noise frequency band for which the confidence level is below a threshold confidence level. According to some such examples, modifying the default noise compensation method may involve reducing a default noise compensation level adjustment for one or more frequency bands.

[0121] In some examples, confidence values between a minimum value and a threshold confidence level (e.g., a maximum confidence level) can correspond to linear interpolations between a minimum applied gain and a "default" level adjustment for noise compensation. In some implementations, a minimum (e.g., zero) confidence score can correspond to a timbre-preserving noise compensation method, and a maximum (e.g., 1.0) confidence score can correspond to an unconstrained or "default" level adjustment for noise compensation. In some examples, the confidence values can be inversely proportional to the timbre-preserving setting. For example, if the minimum confidence value is 0.0 and the maximum confidence value is 1.0, the minimum (e.g., zero) confidence score can correspond to a timbre-preserving setting of 100% or 1.0. In some examples, a timbre-preserving setting of 0.50 can correspond to a confidence value of 0.5. In some such examples, a confidence value of 0.25 can correspond to a timbre-preserving setting of 0.75.

[0122] According to some examples, method 300 may involve receiving a content stream including audio data by a control system and via an interface system. In some such examples, implementing a noise compensation method in block 325 may involve applying the noise compensation method to the audio data to generate noise-compensated audio data. Some such embodiments may involve providing the noise-compensated audio data to one or more audio reproduction transducers of an audio environment by a control system and via an interface system. Some such embodiments may involve rendering the noise-compensated audio data by a control system to generate a rendered audio signal. Some such embodiments may involve providing the rendered audio signal to at least some of a group of audio reproduction transducers of an audio environment by a control system and via an interface system.

[0123] Figure 4A and Figure 4B Additional examples of noise compensation system components are shown. Figure 4C It shows that Figure 4A and Figure 4B According to these examples, the noise compensation system 410 includes a television 411, a television microphone 412 configured to sample the audio environment in which the noise compensation system 410 is present, stereo speakers 413 and 414, and a remote control 417 for the television 411. Although not shown in FIG. Figure 4A and Figure 4B, but in this example, the noise compensation system 410 includes a noise estimator and a noise compensator, which may be the noise estimator and the noise compensator described above. Figure 1A Examples of noise estimator 107 and noise compensator 102 are described.

[0124] In some examples, the noise estimator 107 and the noise compensator 102 can be controlled via a control system, such as the control system of the television 411 (which may be a control system of the television 411 described below). Figure 2 210 ) may be implemented, for example, according to instructions stored on one or more non-transitory storage media. Similarly, depending on the particular implementation, reference may be made to Figures 4A to 5 The described operations may be performed via the control system of the television 411, via the control system of the remote control 417, or via both control systems. Figures 4A to 5 The noise compensation method described can be implemented via a control system of a device other than a television and / or a remote control device, such as a control system of another device with a display (e.g., a laptop computer), a control system of a smart speaker, a control system of a smart hub, a control system of another device with an audio system, etc. In some embodiments, a smart phone (cellular phone) or a smart speaker (e.g., a smart speaker configured to provide virtual assistant functionality) can be configured to perform the reference Figures 4A to 4C is described as being performed by remote control 417. As with the other figures provided herein, Figures 4A to 4C The types, quantities, and arrangements of the elements shown in the figures are provided as examples only. Other embodiments may include more, fewer, and / or different types, quantities, or arrangements of elements, e.g., more loudspeakers and / or more microphones, more or fewer operations, etc. For example, in other embodiments, the arrangement of the elements on the remote control 417 (e.g., the remote microphone 253, the radio transceiver 252B, and / or the infrared (IR) transmitter 251) may be different. In some such examples, the radio transceiver 252B and / or the infrared (IR) transmitter 251 may be present on the front side of the remote control 417 (e.g., on the left). Figure 4B is shown pointing toward the side of the television 411).

[0125] exist Figure 4A In the example shown in FIG, the audio environment 400 in which the noise compensation system 410 is present also includes a listener 416 (assumed to be stationary in this example) and a noise source 415 that is closer to the television 411 than the listener 416. The type and location of the noise source 415 are shown only as an example. In alternative examples, the listener 416 may not be stationary. In some such examples, the listener 416 will be assumed to be in the same location as a remote control 417 or another device that can provide similar functionality, such as a cell phone.

[0126] exist Figure 4A and Figure 4B In the example shown in FIG, remote control 417 is battery powered and incorporates remote microphone 253. In some embodiments, remote control 417 includes remote microphone 253 because remote control 417 is configured to provide voice assistant functionality. To conserve battery life, in some embodiments, remote microphone 253 does not sample ambient noise at all times, nor does remote microphone 253 transmit a continuous stream to television 411. Instead, in some such examples, remote microphone 253 is not always on, but rather "listens" only when remote control 417 receives a corresponding input (e.g., a button press).

[0127] In some examples, the remote control microphone 253 can be used to provide noise level measurements when polled by the television 411 to resolve proximity issues. In some such embodiments, the remote control microphone 253 can respond to a call from the television 411 to the remote control 417 (e.g., Figure 4B The signal from the television 411 may be in response to ambient noise detected by the television microphone 412. Alternatively or additionally, in some examples, the television 411 may poll the remote control 417 at regular intervals to obtain a short-time window ambient noise record. According to some examples, the television 411 may interrupt the polling when the ambient noise decreases. In some alternative examples, the television 411 may be configured to interrupt the polling when the television 411 has sufficient conversion function between the noise level at the location of the remote control 417 and the noise level at the location of the television microphone 412. According to some such embodiments, the television 411 may be configured to resume polling upon receiving an indication that the remote control 417 has moved, for example, upon receiving an inertial sensor signal from the remote control 417 corresponding to the movement. In some embodiments, the level of recording made via the remote control microphone 253 can be used to determine whether the noise estimate made at the television 411 is valid for the listener position (assumed in some examples to correspond to the position of the remote control 417), thereby ensuring that background noise is not over-compensated or under-compensated due to proximity errors.

[0128] According to some examples, based on a polling request from television 411, remote control 417 may communicate with the television 411 across a wireless connection (e.g., from a Figure 4B2B to radio transceiver 252A) sends a short recording of the audio detected by remote control microphone 253 to television 411. The control system of television 411 can be configured to remove the output of loudspeakers 413 and 414 from the recording, for example, by passing the recording through an echo canceller. In some examples, the control system of television 411 can be configured to compare the residual noise recording with the ambient noise detected by television microphone 412 to determine whether the ambient noise from the noise source is louder at television microphone 412 or at the listener's position. In some embodiments, the noise estimate made based on the input from television microphone 412 can be scaled accordingly, for example, based on the ratio of the ambient noise level detected by remote control microphone 253 to the ambient noise level detected by television microphone 412.

[0129] According to some embodiments, the signal transmitted by the infrared (IR) transmitter 251 of the remote control 417 and received by the IR receiver 250 of the television 411 can be used as a synchronization reference, for example, to time-align the echo reference with the remote control's recording for the purpose of echo cancellation. Such an embodiment can solve the clock synchronization problem between the remote control 417 and the television 411 without continuously transmitting the clock signal, which would have an unacceptable impact on battery life.

[0130] Figure 4C A detailed example of one such embodiment is shown in FIG. In this example, time is depicted as the horizontal axis and various operations are illustrated as sections perpendicular to the vertical axis. In this example, Figure 4A and Figure 4B The audio played back by television speakers 413 and 414 is represented as waveform 261.

[0131] According to this example, the television 411 transmits a radio signal 271 to the remote control 417 via the radio transceiver 252A. For example, the radio signal 271 may be transmitted in response to ambient noise detected by the television microphone 412. In this example, the radio signal 271 includes instructions for causing the remote control 417 to record an audio segment via the remote control microphone 253. In some examples, the radio signal 271 may include a start time (e.g., Figure 4C The time T shown in ref ), information used to determine the start time, time interval, etc.

[0132] In this example, the remote control 417 records the audio segment time interval T. rec During this time, the signal received by the remote control microphone 253 is recorded as an audio segment 272. According to this example, the remote control 417 sends a signal to the television 411 indicating that the audio segment time interval T has been recorded. recHere, the signal 265 indicates the recorded audio segment time interval T rec At time T ref 265 stops being transmitted. In this example, the remote control 417 transmits the signal 265 via the IR transmitter 251. Thus, the television 411 can identify the time interval T at which the audio segment was recorded. rec The time interval during which the television speakers 413 and 414 are reproducing the content stream audio segment 269.

[0133] In this example, remote control 417 then sends signal 266 including the recorded audio segment to television 411. According to this embodiment, the control system of television 411 performs an echo cancellation process based on the recorded audio segment and the content stream audio segment 269 to obtain an ambient noise signal 270 at the location of remote control 417, which in this example is assumed to correspond to the location of listener 416. In some such embodiments, the control system of television 411 is configured to perform a noise compensation method on the audio data to be reproduced by television loudspeakers 413 and 414 based at least in part on ambient noise signal 270 to produce noise-compensated audio data.

[0134] Figure 5 500 is a flowchart outlining an example of the disclosed method. As with other methods described herein, the blocks of method 500 do not have to be executed in the order indicated. In addition, such a method may include more or fewer blocks than shown and / or described.

[0135] Method 500 may be performed as follows: Figure 2 2 and described above. In some examples, the blocks of method 500 can be performed by one or more devices within the audio environment, for example, an audio system controller or another component of the audio system, such as a smart speaker, a television, a television control module, a smart speaker, a mobile device, etc. In some embodiments, the audio environment can include one or more rooms of a home environment. In other examples, the audio environment can be another type of environment, such as an office environment, a car environment, a train environment, a street or sidewalk environment, a park environment, etc. However, in alternative embodiments, at least some blocks of method 500 can be performed by a device (such as a server) that implements a cloud-based service.

[0136] In this embodiment, block 505 involves controlling the system by a first device and receiving a content stream including content audio data via a first interface system of the first device in the audio environment. According to some examples, the first device may be a television or a television control module. In some such examples, the content stream may also include content video data corresponding to the content audio data. However, in other examples, the first device may be another type of device, such as a laptop, a smart speaker, a sound bar, etc.

[0137] In this example, block 510 involves receiving, by the first device control system and via the first interface system, a first microphone signal from a first device microphone system of the first device. The first device microphone system may include one or more microphones. According to some examples where the first device is a television or a television control module, a first microphone signal may be received from one or more microphones in, on, or near the television (as described above with reference to FIG. Figure 4A and Figure 4B According to this embodiment, block 515 involves controlling the system by the first device and detecting ambient noise from a noise source location in or near the audio environment based at least in part on the first microphone signal.

[0138] According to this example, block 520 involves causing, by the first device control system, a first wireless signal to be transmitted from the first device via the first interface system to a second device in the audio environment. In this example, the first wireless signal includes instructions for causing the second device to record an audio segment via the second device microphone system. In some embodiments, the second device may be a remote control device, a smart phone, or a smart speaker. According to some examples, the first wireless signal may be transmitted via radio waves or microwaves. In some examples, block 520 may involve transmitting signal 271, as described above with reference to Figure 4C According to this example, the first wireless signal is responsive to detecting ambient noise in block 515. According to some examples, the first wireless signal may be responsive to determining that the detected ambient noise is greater than or equal to an ambient noise threshold level.

[0139] In some instances, the first wireless signal may include the second device audio recording start time or information for determining the second device audio recording start time. In some examples, the information for determining the second device audio recording start time may include or may be instructions for waiting until frequency hopping occurs when the first wireless signal is transmitted via a frequency hopping system (e.g., Bluetooth). In some examples, the information for determining the second device audio recording start time may include or may be instructions for waiting until a time slot is available when the first wireless signal is transmitted via a time division multiplexing wireless system. In some examples, the first wireless signal may indicate the second device audio recording time interval.

[0140] According to this example, block 525 involves controlling the system by the first device and receiving a second wireless signal from the second device via the first interface system. According to some examples, the second wireless signal may be sent via infrared waves. In some examples, block 525 may involve receiving signal 265, as described above with reference to Figure 4C As described. In some examples, the second wireless signal may indicate a start time of the second device audio recording. In some examples, the second wireless signal may indicate a time interval of the second device audio recording. According to some examples, the second wireless signal (or a subsequent signal from the second device) may indicate an end time of the second device audio recording. In some such examples, method 500 may involve controlling the system by the first device and receiving a fourth wireless signal from the second device via the first interface system, the fourth wireless signal indicating an end time of the second device audio recording.

[0141] In this example, block 530 involves determining, by the first device control system, a content stream audio segment time interval for the content stream audio segment. In some examples, block 530 may involve determining a time interval for the content stream audio segment 269, as described above with reference to Figure 4C In some examples, the first device control system content stream may be configured to determine the audio segment time interval based on the second device audio recording start time and the second device audio recording end time, or based on the second device audio recording start time and the second device audio recording time interval. In some examples involving receiving a fourth wireless signal from the second device indicating the second device audio recording end time, method 500 may involve determining the content stream audio segment end time based on the second device audio recording end time.

[0142] According to this example, block 535 involves controlling the system by the first device and receiving a third wireless signal from the second device via the first interface system, the third wireless signal including the recorded audio segment captured via the second device microphone. In some examples, block 535 may involve receiving signal 266, as described above with reference to Figure 4C described.

[0143] In this example, block 540 involves determining, by the first device control system, a second device ambient noise signal at the second device location based at least in part on the recorded audio segment and the content stream audio segment. In some examples, block 540 can involve performing an echo cancellation process based on the recorded audio segment and the content stream audio segment 269 to obtain the ambient noise signal 270 at the location of the remote control 417, as described above with reference to Figure 4C described.

[0144] According to this example, block 545 involves implementing, by the first device control system, a noise compensation method on the content audio data based at least in part on the second device ambient noise signal to produce noise-compensated audio data. In some examples, method 500 may involve receiving, by the first device control system and via the first interface system, a second microphone signal from the first device microphone system during the second device audio recording time interval. Some such examples may involve detecting, by the first device control system and based at least in part on the first microphone signal, a first device ambient noise signal corresponding to ambient noise from a noise source location. In such examples, the noise compensation method may be based at least in part on the first device ambient noise signal.

[0145] According to some such examples, the noise compensation method can be based at least in part on a comparison of a noise signal around the first device and a noise signal around the second device. In some examples, the noise compensation method can be based at least in part on a ratio of the noise signal around the first device to the noise signal around the second device.

[0146] Some examples may involve providing (e.g., by the first device controlling the system and via the first interface system) noise-compensated audio data to one or more audio reproduction transducers of the audio environment. Some examples may involve rendering (e.g., by the first device controlling the system) the noise-compensated audio data to produce a rendered audio signal. Some such examples may involve providing (e.g., by the first device controlling the system and via the first interface system) the rendered audio signal to at least some of a set of audio reproduction transducers of the audio environment. In some such examples, at least one of the reproduction transducers of the audio environment may be present in the first device.

[0147] Figure 6 An additional example of a noise compensation system is shown. In this example, Figure 6 An example of a noise compensation system with three microphones is shown that allows the control system to determine the location of the noise source. Figure 6In the example shown in FIG, noise compensation system 710 includes television 711 and television microphones 702a, 702b, and 702c. In some alternative examples, noise compensation system 710 may include a remote control for television 711, which in some instances may be configured to function like remote control 417. Figure 6 The noise compensation system includes a noise estimator and a noise compensator, which are not shown in FIG. Figure 1A Examples of noise estimator 107 and noise compensator 102 are described.

[0148] In some examples, the noise estimator 107 and the noise compensator 102 can be controlled via a control system, such as the control system of the television 611 (which may be a control system of the television 611 described below). Figure 2 210 ) may be implemented, for example, according to instructions stored on one or more non-transitory storage media. Similarly, depending on the particular implementation, reference may be made to Figures 6 to 7B The operations described may be performed via the control system of the television 611, via the control system of the remote controller, or via both control systems. Figures 6 to 7B The described noise compensation method can be implemented via a control system of a device other than a television and / or a remote control device, such as a control system of another device with a display (e.g., a laptop computer), a control system of a smart speaker, a control system of a smart hub, a control system of another device of an audio system, etc.

[0149] exist Figure 6 In the example shown in FIG, the audio environment in which noise compensation system 710 is present also includes a listener 616 (assumed to be stationary in this example) and a noise source 615. In some examples, the position of listener 616 can be assumed to be the same as or closely adjacent to the position of the remote control. In this instance, noise source 615 is closer to listener 616 than television 611. The type and location of noise source 615 are shown only as an example.

[0150] As with the other figures presented in this article, Figures 6 to 7A The types, quantities, and arrangements of the elements shown in the drawings are provided as examples only. Other implementations may include more, fewer, and / or different types, quantities, or arrangements of elements, for example, more loudspeakers and / or more microphones, more or fewer operations, etc.

[0151] Figure 6 An example of acoustic propagation paths 707a, 707b, and 707c from a noise source 615 to a microphone 702 is shown. In this example, the acoustic propagation paths 707a, 707b, and 707c have different lengths and therefore arrive at each microphone at different times. Figure 6In the example shown in , no synchronization is required between the multiple devices because microphones 702a, 702b, and 702c are part of television 711 and are controlled by the same control system.

[0152] According to some examples, a cross-correlation function of the recorded ambient noise from microphones 702a, 702b, and 702c can be calculated to determine the arrival time difference between the microphones. The path length difference is the time difference (seconds) multiplied by the speed of sound (meters / second). Based on the path length difference, the distance from the listener 616 to the TV 711, and the known distances between the microphones 702a, 702b, and 702c, the position of the noise source 615 can be solved. In some examples, the position of the noise source 615 can be calculated using a two-dimensional (2D) hyperbolic position location algorithm, such as one of the methods described in Chapter 1.21, 1.22, 2.1, or 2.2 of Dalskov, D., Locating Acoustic Sources with Multilateration - Applied to Stationary and Moving Sources (Aalborg University, June 4, 2014), which is hereby incorporated by reference. Reference is made below to Figure 7A and Figure 7B A specific example of an alternative solution is described.

[0153] Figure 7A Is indicated by Figure 6 An example of an image of a signal received by a microphone is shown in FIG. In this example, Figure 7A An example correlation analysis of three microphones is shown to determine the time difference of arrival (TDOA) of microphones 702a and 702c of a noise source 615 relative to the center microphone 702b. According to this example, Figure 7A The elements are as follows:

[0154] 712a represents the cross-correlation between microphone 702a and reference microphone 702b;

[0155] 712b represents the autocorrelation of reference microphone 702b;

[0156] 712c represents the cross-correlation between microphone 702c and reference microphone 702b;

[0157] 713a is the peak in the cross-correlation that determines the TDOA of microphone 702a relative to reference microphone 702b. In this example, it can be seen that the sound arrives at microphone 702a first and then at reference microphone 702b, resulting in a negative TDOA for microphone 702a.

[0158] 713b is the peak in the autocorrelation of reference microphone 702b. In this example, time 0 is defined as the location of this peak. In some alternative embodiments, the autocorrelation function 712b may be deconvolved with the cross-correlation functions 712a and 712c before estimating TDOA to form a sharper peak.

[0159] 713c is the peak in the cross-correlation that can determine the TDOA of microphone 702c relative to reference microphone 702b. In this example, it can be seen that the sound arrives at reference microphone 702b first before arriving at microphone 702c, resulting in a positive TDOA for microphone 702c.

[0160] 714a is a visual representation of the TDOA of microphone 702a relative to reference microphone 702b. Mathematically, TDOA 714a would be considered a negative quantity in this example because the sound arrives at microphone 702a before arriving at microphone 702b; and

[0161] 714b is a visual representation of the TDOA of microphone 702c relative to reference microphone 702b. This TDOA will be a positive quantity in this example.

[0162] Figure 7B Shows the audio environment in different locations Figure 6 In this example, the noise source has been redrawn Figure 6 The arrangement shown in FIG. 6 is to emphasize the geometric nature of the problem and to mark the length of each side of each triangle. In this example, the noise source 615 is shown as coming from the right side of the figure rather than the left side, as shown in FIG. Figure 6 and Figure 7A This makes the x-coordinate 720a of the noise source a positive quantity to help clearly define the coordinate system.

[0163] exist Figure 7B In the example shown, the elements are as follows:

[0164] 615 represents the noise source to be located;

[0165] 702a to 702c indicate Figure 6 Here, the reference microphone 702b is shown as the origin of the two-dimensional Cartesian coordinate system;

[0166] 720a represents the x-coordinate (in meters) of the noise source 615 relative to the origin located at the center of the reference microphone 702b;

[0167] 720b represents the y-coordinate (in meters) of the noise source 615 relative to the origin located at the center of the reference microphone 702b;

[0168] 721a represents the distance between microphone 702a and microphone 702b. In this example, microphone 702a is positioned on the TV, d meters to the left of reference microphone 702b. In one example, d = 0.4m;

[0169] 721b represents the distance between microphone 702b and microphone 702c. In this example, microphone 702a is positioned on the TV, d meters to the right of reference microphone 702b;

[0170] 722 represents the noise source 615 projected onto the x-axis of the Cartesian coordinate system;

[0171] 707a to 707c represent the acoustic path length (in meters) from the noise source 615 to each of the microphones 702a to 702c;

[0172] 708b corresponds to the symbol r, which in this example is defined to mean the distance in meters from the noise source 615 to the reference microphone 702b;

[0173] 708a corresponds to the sum of the symbols r+a. In this example, the symbol a is defined to mean the difference in path length between 707a and 707b, such that the length of acoustic path 707a is r+a. The TDOA of microphone 702a relative to microphone 702b (see Figure 7A (714a in that example is negative, but positive in this example) calculate the acoustic path length r+a by multiplying the TDOA by the speed of sound in the medium. For example, if the TDOA 714a is +0.0007s and the speed of sound is 343 m / s, then a=0.2401 m;

[0174] 708c corresponds to the sum of the symbols r+b. In this example, the symbol b is defined to mean the difference in path length between 707c and 707b, such that the length of acoustic path 707c is r+b. The TDOA of microphone 702c relative to microphone 702b (see Figure 7A The acoustic path length r+b is calculated by multiplying TDOA by the speed of sound in the medium (in 714c, which is positive in that example but negative in this example). For example, if TDOA 714c is -0.0006 s and the speed of sound is 343 m / s, then b = -0.2058 m. In some embodiments, the control system can be configured to determine a more accurate speed of sound for the audio environment based on input from a temperature sensor.

[0175] Now write the Pythagoras' Theorem for triangle (702b, 615, 722):

[0176] r 2 =x 2 +y 2 ...Equation 2

[0177] The Pythagorean theorem for triangle (702a, 615, 722) can be written as follows:

[0178] (r+a) 2 =(x+d) 2 +y 2 ...Equation 3

[0179] The Pythagorean theorem for triangle (702c, 615, 722) can be written as follows:

[0180] (r+b) 2 =(xd) 2 +y 2 ...Equation 4

[0181] Equations 2, 3, and 4 together form a system of three simultaneous equations consisting of the unknowns r, x, and y. Of particular interest is r, the distance in meters from the noise source to the reference microphone 702b.

[0182] This system of equations can be solved to find r as follows:

[0183] r=-(a 2 +b 2 -2d 2 ) / (2(a+b))...Equation 5

[0184] For the example values given above:

[0185] a=0.2401m, b=-0.2058m, d=0.4m,

[0186] It can be concluded that r = 3.206 m. Therefore, in this example, the noise source 615 is located approximately 3.2 m from the reference microphone 702 b.

[0187] In addition to estimating the location of noise sources, some embodiments may also involve determining or estimating the listener location. Figure 6 Depending on the particular implementation, the distance from the listener 616 to the television 611 can be estimated or determined in different ways. According to some examples, the distance from the listener 616 to the television 611 can be determined by the listener 616 or by another user during the initial setup of the television 611. In other examples, the distance can be estimated or determined based on input from one or more sensors (e.g., based on input from one or more cameras, based on input from additional microphones, or based on input from the sensors described above). Figure 2The position of the listener 161 and / or the distance from the listener 616 to the television 611 may be determined using input from the sensor system 230 and other sensors described herein. In other examples, in the absence of user input or sensor input, the distance from the listener 616 to the television 611 may be determined based on a default distance, which may be an average distance from a typical listener to the television. In some examples, it may be assumed that the listener is within a certain angle from the normal to the television screen, e.g., within 10 degrees, within 15 degrees, within 20 degrees, within 25 degrees, within 30 degrees, etc.

[0188] According to some embodiments, noise compensation can be based at least in part on the determined or estimated listener position and the determined or estimated noise source position. For example, by knowing where listener 616 is (or assuming where listener 616 is relative to television 711) and knowing the position of noise source 615 and the corresponding noise level at television 711, a propagation loss model can be used to calculate a noise estimate for the position of listener 616. This predicted noise compensation value for the listener's position can be used directly by the noise compensation system.

[0189] In some alternative embodiments, the predicted noise level at the listener position can be further modified to include a confidence value. For example, if the noise source is relatively far from the listener position (or far from multiple most likely listener positions and the predicted noise estimate does not vary greatly between the most likely listener positions), the noise estimate will have a high confidence. Otherwise, the noise estimate may have a lower confidence. The list of possible listener positions may change depending on the context of the system. In addition, according to some examples, if there may be multiple microphones measuring noise levels at various locations in the audio environment, the noise estimate confidence may be further increased. Compared to the case where the noise levels measured at various locations are inconsistent with the propagation loss model, the case where the noise levels measured at various locations in the audio environment are all consistent with the propagation loss model can provide a higher confidence in the noise estimate.

[0190] If the noise compensation system has a high confidence level in the noise estimate at the listener's position, then in some embodiments, the noise compensation system can be configured to implement an unconstrained noise compensation method. Alternatively, if the noise compensation system has a low confidence level in the noise estimate at the listener's position, then the noise compensation system can implement a more constrained noise compensation method.

[0191] Figure 8 An example of a floor plan of an audio environment, which in this example is a living space, is shown. As with the other figures provided herein, Figure 8The types, quantities, and arrangements of the elements shown in the drawings are provided as examples only. Other implementations may include more, fewer, and / or different types, quantities, or arrangements of elements.

[0192] According to this example, the environment 800 includes a living room 810 at the upper left, a kitchen 815 at the lower center, and a bedroom 822 at the lower right. The boxes and circles distributed across the living space represent a set of speakers 805a to 805h, at least some of which may be smart speakers in some embodiments, placed in locations convenient for the space but not following any standard prescribed layout (arbitrarily placed). In some examples, a television 830 may be configured to at least partially implement one or more disclosed embodiments. In this example, the environment 800 includes cameras 811a to 811e distributed throughout the environment. In some embodiments, one or more smart audio devices in the environment 800 may also include one or more cameras. The one or more smart audio devices may be single-purpose audio devices or virtual assistants. In some such examples, one or more cameras of the optional sensor system 130 may be present in or on the television 830, in a mobile phone, or in a smart speaker (such as one or more of speakers 805b, 805d, 805e, or 805h). Although cameras 811a through 811e are not shown in every depiction of environments 800 presented in this disclosure, in some implementations, each of environments 800 may still include one or more cameras.

[0193] Some aspects of the present disclosure include a system or device configured (e.g., programmed) to perform one or more examples of the disclosed methods, and a tangible computer-readable medium (e.g., a disk) storing code for implementing one or more examples of the disclosed methods or their steps. For example, some disclosed systems may be or include a programmable general-purpose processor, a digital signal processor, or a microprocessor that is programmed and / or otherwise configured with software or firmware to perform any of a variety of operations on data, including embodiments of the disclosed methods or their steps. Such a general-purpose processor may be or include a computer system that includes an input device, a memory, and a processing subsystem that is programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or their steps) in response to data asserted thereto.

[0194] Some embodiments can be implemented as a configurable (e.g., programmable) digital signal processor (DSP), which is configured (e.g., programmed and otherwise configured) to perform the required processing on (multiple) audio signals, including the execution of one or more examples of the disclosed method. Alternatively, the embodiment of the disclosed system (or its elements) can be implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which can include an input device and a memory), which is programmed with software or firmware and / or otherwise configured to perform any of the various operations, including one or more examples of the disclosed method. Alternatively, the elements of some embodiments of the present invention system are implemented as a general-purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed method, and the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). The general-purpose processor configured to perform one or more examples of the disclosed method can be coupled to an input device (e.g., a mouse and / or keyboard), a memory, and a display device.

[0195] Another aspect of the present disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) storing code for performing one or more examples of the disclosed method or its steps (e.g., an encoder executable to perform one or more examples of the disclosed method or its steps).

[0196] Although specific embodiments of the present disclosure and applications of the present disclosure have been described herein, it will be apparent to those skilled in the art that many changes may be made to the embodiments and applications described herein without departing from the scope of the present disclosure as described and claimed herein. It should be understood that although certain forms of the present disclosure have been shown and described, the present disclosure is not limited to the specific embodiments described and shown or the specific methods described.

Claims

1. A noise compensation method, comprising: receiving, by the control system and via the interface system, a microphone signal corresponding to ambient noise from a noise source location in or near the audio environment; determining or estimating, by the control system, a listener position in the audio environment; estimating, by the control system, at least one critical distance, the critical distance being a distance from the location of the noise source at which a sound pressure of directly propagated ambient noise from the noise source is equal to a diffuse field sound pressure of the ambient noise from the noise source; estimating whether the listener position is within the at least one critical distance from the noise source position; as well as A noise compensation method is performed on the ambient noise based at least in part on at least one estimate of whether the listener position is within the at least one critical distance from the noise source position.

2. The noise compensation method according to claim 1, further comprising: controlling, via the control system, an audio reproduction transducer system in an audio environment to reproduce one or more room calibration sounds, the audio reproduction transducer system comprising one or more audio reproduction transducers; receiving, by the control system and via the interface system, a microphone signal corresponding to a response of the audio environment to the one or more room calibration sounds; determining, by the control system and based on the microphone signal, a reverberation time for each of a plurality of frequencies; as well as determining or estimating an audio environment volume of the audio environment, wherein estimating the at least one critical distance involves calculating a plurality of estimated frequency-based critical distances based at least in part on the plurality of frequency-related reverberation times and the audio environment volume, each estimated frequency-based critical distance of the plurality of estimated frequency-based critical distances corresponding to a frequency of the plurality of frequencies.

3. The noise compensation method according to claim 2, wherein: Estimating whether the listener position is within the at least one critical distance involves estimating whether the listener position is within each frequency-based critical distance of the plurality of frequency-based critical distances. 4 . The noise compensation method of claim 3 , further comprising transforming the microphone signal corresponding to the ambient noise from the time domain into the frequency domain, and determining a frequency band ambient noise level estimate for each of a plurality of ambient noise frequency bands. 5 . The noise compensation method of claim 4 , further comprising determining a frequency-based confidence level for each of the ambient noise level estimates.

6. The noise compensation method according to claim 5, wherein: Each frequency-based confidence level corresponds to an estimate of whether the listener position is within each frequency-based critical distance.

7. The noise compensation method according to claim 5, wherein: Each frequency-based confidence level is inversely proportional to each frequency-based critical distance.

8. The noise compensation method according to claim 5, wherein: Implementing the noise compensation method involves implementing a frequency-based noise compensation method based on the frequency-based confidence level for each ambient noise frequency band.

9. The noise compensation method according to claim 8, wherein: The frequency-based noise compensation method involves applying a default noise compensation method for each ambient noise frequency band for which the confidence level is at or above a threshold confidence level.

10. The noise compensation method according to claim 8, wherein: The frequency-based noise compensation method involves modifying a default noise compensation method for each ambient noise frequency band for which the confidence level is below a threshold confidence level.

11. The noise compensation method according to claim 10, wherein: Modifying the default noise compensation method involves reducing a default noise compensation level adjustment.

12. The noise compensation method according to claim 2, wherein: The one or more room calibration sounds are embedded in content audio data received by the control system.

13. The noise compensation method according to any one of claims 1 to 12, further comprising receiving, by the control system and via the interface system, a content stream comprising audio data, wherein Implementing the noise compensation method involves applying the noise compensation method to the audio data to produce noise-compensated audio data.

14. The noise compensation method of claim 13, further comprising providing, by the control system and via the interface system, the noise compensated audio data to one or more audio reproduction transducers of the audio environment.

15. The noise compensation method according to claim 13, further comprising: rendering, by the control system, the noise-compensated audio data to produce a rendered audio signal; as well as The rendered audio signal is provided by the control system and via the interface system to at least some of a set of audio reproduction transducers of the audio environment. 16 . An apparatus for noise compensation, the apparatus comprising components configured to implement the noise compensation method according to claim 1 .

17. One or more non-transitory media having software stored thereon, the software comprising instructions for controlling one or more devices to implement the noise compensation method according to any one of claims 1 to 15.

18. A system for noise compensation, the system comprising components configured to implement the noise compensation method according to any one of claims 1 to 15.

19. A computer program product comprising computer-executable instructions for executing the method according to any one of claims 1 to 15 by one or more processors.

Citation Information

Patent Citations

  • Speech Signal Enhancement Using Visual Information

    US20140337016A1

  • Method, computer readable storage medium, and apparatus for multichannel audio playback adaptation for multiple listening positions

    US20170245083A1