Pervasive Acoustic Mapping
The system addresses the challenge of inefficient acoustic scene estimation in audio systems by using inaudible DSSS signals for automated acoustic mapping, enabling flexible device placement and adaptive calibration.
Patent Information
- Application Number
- JP2023533816
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-28
- Filing Date
- 2021-12-02
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2041-12-02
AI Technical Summary
Existing audio systems lack efficient methods for estimating acoustic scene metrics, such as audio device audibility, and require intrusive calibration procedures that hinder widespread adoption and fail to adapt to changes in the acoustic environment.
A system that generates and injects inaudible direct sequence spread spectrum (DSSS) signals into audio content, allowing audio devices to self-organize and calibrate by extracting these signals from microphone inputs to estimate acoustic scene metrics, enabling automated and adaptive acoustic mapping.
Enables flexible and automated acoustic mapping that adapts to changes in the audio environment, improving user experience by allowing devices to be placed anywhere without manual intervention and providing accurate spatial reproduction and voice interaction.
Smart Images

Figure 0007815249000025 
Figure 0007815249000026 
Figure 0007815249000027
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of Spanish Patent Application No. P202031212, filed December 3, 2020; U.S. Provisional Patent Application No. 63 / 120,963, filed December 3, 2020; U.S. Provisional Patent Application No. 63 / 120,887, filed December 3, 2020; U.S. Provisional Patent Application No. 63 / 121,007, filed December 3, 2020; U.S. Provisional Patent Application No. 63 / 121,085, filed on March 3, 2021; U.S. Provisional Patent Application No. 63 / 155,369, filed on March 2, 2021; U.S. Provisional Patent Application No. 63 / 201,561, filed on May 4, 2021; Spanish Patent Application No. P202130458, filed on May 20, 2021; U.S. Provisional Patent Application No. P202130458, filed on July 21, 2021 This application claims benefit of priority to U.S. Provisional Patent Application No. 63 / 203,403; U.S. Provisional Patent Application No. 63 / 224,778, filed July 22, 2021; Spanish Patent Application No. P202130724, filed July 26, 2021; U.S. Provisional Patent Application No. 63 / 260,528, filed August 24, 2021; U.S. Provisional Patent Application No. 63 / 260,529, filed August 24, 2021; U.S. Provisional Patent Application No. 63 / 260,953, filed September 7, 2021; U.S. Provisional Patent Application No. 63 / 260,954, filed September 7, 2021; and U.S. Provisional Patent Application No. 63 / 261,769, filed September 28, 2021, all of which are incorporated herein by reference.
[0002] Technical Field TECHNICAL FIELD The present disclosure relates to audio processing systems and methods. [Background technology]
[0003] Audio devices and systems are widely deployed, and while existing systems and methods for estimating acoustic scene metrics (e.g., audio device audibility) are known, improved systems and methods would be desirable.
[0004] Notation and Name Throughout this disclosure, including the claims, the terms "speaker," "loudspeaker," and "audio reproduction transducer" are used interchangeably to refer to any sound-emitting transducer (or collection of transducers). A typical set of headphones includes two speakers. A speaker may be implemented to include multiple transducers (e.g., woofers and tweeters) that may be driven by a single common speaker feed or by multiple speaker feeds. In some examples, the speaker feeds may receive different processing in different circuit branches coupled to different transducers.
[0005] Throughout this disclosure, including the claims, the phrase performing an operation "on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to the signal or data) is used broadly to refer to performing the operation directly on the signal or data, or to performing the operation on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or preprocessing before performing the operation).
[0006] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M of the inputs and the other XM inputs are received from external sources) may also be referred to as a decoder system.
[0007] Throughout this disclosure, including the claims, the term "processor" is used broadly to refer to a system or device that is programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, video, or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipeline processing on audio or other audio data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.
[0008] Throughout this disclosure, including the claims, the terms "couple" or "coupled" are used to mean a direct or indirect connection. Thus, when a first device couples to a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections.
[0009] As used herein, a "smart device" is an electronic device typically configured to communicate with one or more other devices (or networks) via various wireless protocols, such as Bluetooth, Zigbee, near-field communications, Wi-Fi, light fidelity (Li-Fi), 3G, 4G, or 5G, and capable of operating interactively and / or autonomously to some degree. Some notable types of smart devices are smartphones, smart cars, smart thermostats, smart doorbells, smart locks, smart refrigerators, phablets and tablets, smart watches, smart bands, smart keychains, and smart audio devices. The term "smart device" may also refer to devices that exhibit some characteristics of ubiquitous computing, such as artificial intelligence.
[0010] As used herein, the phrase "smart audio device" refers to a smart device that is either a single-purpose audio device or a multi-purpose audio device (e.g., an audio device that implements at least some aspects of virtual assistant functionality). A single-purpose audio device is a device (e.g., a television (TV)) that includes or is coupled to at least one microphone (and optionally includes or is coupled to at least one speaker and / or at least one camera) and is designed largely or primarily to achieve a single purpose. For example, while televisions are typically capable of (and are considered capable of) playing audio from program material, modern televisions most often run some kind of operating system on which applications, including television viewing applications, run locally. In this sense, a single-purpose audio device with a speaker and microphone is often configured to run local applications and / or services that directly use the speaker and microphone. Several single-purpose audio devices can be configured to group together to achieve audio playback over a zone or user-configured area.
[0011] One common type of multipurpose audio device is an audio device that implements at least some aspects of virtual assistant functionality, although other aspects of the virtual assistant functionality may be implemented by one or more other devices, such as one or more servers with which the multipurpose audio device is configured to communicate. Such multipurpose audio devices are sometimes referred to herein as "virtual assistants." A virtual assistant is a device (e.g., a smart speaker or voice-assistant-integrated device) that includes or is coupled to at least one microphone (and, optionally, includes or is coupled to at least one speaker and / or at least one camera). In some examples, a virtual assistant may provide the ability to utilize multiple devices (different from the virtual assistant) for applications that are in some way cloud-enabled or otherwise not entirely implemented within or on the virtual assistant itself. In other words, at least some aspects of the virtual assistant functionality, e.g., speech recognition functionality, may be implemented (at least in part) by one or more servers or other devices with which the virtual assistant can communicate over a network such as the Internet. Virtual assistants may sometimes cooperate, for example, in a discrete, conditionally defined manner. For example, two or more virtual assistants can cooperate in the sense that one of them, e.g., the virtual assistant that is most confident that it heard the wake word, will respond to that word. Connected virtual assistants can, in some implementations, form a kind of constellation, which may be managed by one main application that may be (or may implement) a virtual assistant.
[0012] Here, "wake word" is used broadly to mean any sound (e.g., a word spoken by a human being, or some other sound) that the smart audio device is configured to wake up in response to detecting ("listening") for that sound (using at least one microphone included in or coupled to the smart audio device, or at least one other microphone). In this context, "awakening" refers to the device entering a state in which it waits for (i.e., listens for) a voice command. In some instances, what may be referred to herein as a "wake word" may include multiple words, e.g., a phrase.
[0013] Here, the term "wake word detector" refers to a device (or software including instructions for configuring a device) configured to continuously look for alignment between real-time audio (e.g., speech) features and a trained model. Typically, a wake word event is triggered whenever the wake word detector determines that the probability that the wake word has been detected exceeds a predetermined threshold. For example, the threshold may be a predetermined threshold adjusted to provide a reasonable compromise between false accept and false reject rates. Following a wake word event, the device may enter a state (which may be referred to as an "awake" or "attentive" state) in which it listens for commands and passes received commands to a larger, more computationally intensive recognizer.
[0014] As used herein, the terms "program stream" and "content stream" refer to a collection of one or more audio signals, and possibly video signals, at least portions of which are intended to be listened to together. Examples include selections of music, movie soundtracks, movies, television programs, audio portions of television programs, podcasts, live voice calls, synthesized voice responses from smart assistants, etc. In some cases, a content stream may contain multiple versions of at least a portion of an audio signal, e.g., the same dialogue in multiple languages. In such cases, only one version of the audio data or portion thereof (e.g., the version corresponding to a single language) is intended to be played at a time. Summary of the Invention [Means for solving the problem]
[0015] At least some aspects of the present disclosure may be implemented via one or more audio processing methods. In some cases, the methods may be implemented, at least in part, by a control system and / or via instructions (e.g., software) stored on one or more non-transitory media. Some methods may involve causing a first audio device of an audio environment, by the control system, to generate a first calibration signal and inserting, by the control system, the first calibration signal into a first audio playback signal corresponding to a first content stream to generate a first modified audio playback signal for the first audio device. Some such methods may involve causing, by the control system, the first audio device to play the first modified audio playback signal to generate the first audio device playback sound.
[0016] Some such methods may involve causing a second audio device in the audio environment to generate a second calibration signal by a control system; causing the control system to insert the second calibration signal into a second content stream to generate a second modified audio playback signal for the second audio device; and causing the control system to play the second modified audio playback signal to generate second audio device playback sound.
[0017] Some such methods may involve causing at least one microphone in the audio environment to detect at least a first audio device and a second audio device and generate microphone signals corresponding to the at least a first audio device and a second audio device. Some such methods may involve causing the control system to extract a first calibration signal and a second calibration signal from the microphone signals. Some such methods may involve causing the control system to estimate at least one acoustic scene metric based at least in part on the first calibration signal and the second calibration signal.
[0018] In some implementations, the control system may be a master device control system.
[0019] In some examples, the first calibration signal may correspond to a first inaudible component of sound reproduced by a first audio device, and the second calibration signal may correspond to a second inaudible component of sound reproduced by a second audio device. According to some examples, the first calibration signal may be or include a first DSSS signal, and the second calibration signal may be or include a second DSSS signal.
[0020] Some methods may involve causing a control system to insert a first gap in a first frequency range of a first audio playback signal or a first modified audio playback signal during a first time interval of a first content stream. The first gap may be or include an attenuation of the first audio playback signal in the first frequency range. In some such examples, the first modified audio playback signal and the first audio device playback sound may include the first gap.
[0021] Some methods may involve causing a control system to insert a first gap within a first frequency range of the second audio playback signal or the second modified audio playback signal during a first time interval. In some such examples, the second modified audio playback signal and the second audio device playback sound may include the first gap.
[0022] Some methods may involve causing a control system to extract audio data from a microphone signal within at least a first frequency range to generate extracted audio data. Some such methods may involve causing the control system to estimate at least one acoustic scene metric based at least in part on the extracted audio data.
[0023] Some methods may involve controlling gap insertion and calibration signal generation such that the calibration signal does not correspond to a gap time interval or a gap frequency range. Some methods may involve controlling gap insertion and calibration signal generation based at least in part on a time since noise was estimated in at least one frequency band. Some methods may involve controlling gap insertion and calibration signal generation based at least in part on a signal-to-noise ratio of a calibration signal of at least one audio device in at least one frequency band.
[0024] Some methods may involve causing a target audio device to play an unmodified audio playback signal of the target device content stream to generate a target audio device playback sound. Some such methods may involve causing a control system to estimate target audio device audibility and / or target audio device location based at least in part on the extracted audio data. In some such examples, the unmodified audio playback signal does not include the first gap. According to some such examples, the microphone signal may correspond to the target audio device playback sound. In some instances, the unmodified audio playback signal does not include a gap inserted in any frequency range.
[0025] In some examples, the at least one acoustic scene metric includes time of flight, time of arrival, direction of arrival, range, audio device audibility, audio device impulse response, angle between audio devices, audio device position, audio environment noise, signal-to-noise ratio, or a combination thereof. According to some implementations, having the at least one acoustic scene metric estimated may involve estimating the at least one acoustic scene metric. In some implementations, having the at least one acoustic scene metric estimated may involve having another device estimate the at least one acoustic scene metric. Some examples may involve controlling one or more aspects of audio device playback based at least in part on the at least one acoustic scene metric.
[0026] According to some implementations, a first content stream component of sound reproduced by a first audio device may cause perceptual masking of a first calibration signal component of sound reproduced by the first audio device. In some such implementations, a second content stream component of sound reproduced by a second audio device may cause perceptual masking of a second calibration signal component of sound reproduced by the second audio device.
[0027] Some examples may involve causing third through Nth audio devices of the audio environment to generate third through Nth calibration signals, by the control system. Some such examples may involve causing third through Nth calibration signals, by the control system, to be inserted into third through Nth content streams to generate third through Nth modified audio playback signals for the third through Nth audio devices. Some such examples may involve causing third through Nth audio devices, by the control system, to play corresponding instances of the third through Nth modified audio playback signals to generate third through Nth instances of audio device playback sounds.
[0028] Some such examples may involve causing at least one microphone of each of first through Nth audio devices to detect first through Nth instances of audio device reproduced sound and generate microphone signals corresponding to the first through Nth instances of audio device reproduced sound. In some instances, the first through Nth instances of audio device reproduced sound may include a first audio device reproduced sound, a second audio device reproduced sound, and third through Nth instances of audio device reproduced sound. Some such examples may involve causing the control system to extract first through Nth calibration signals from the microphone signals. In some implementations, at least one acoustic scene metric may be estimated based at least in part on the first through Nth calibration signals.
[0029] Some examples may involve determining one or more calibration signal parameters for a plurality of audio devices in an audio environment. In some instances, the one or more calibration signal parameters may be usable for generating a calibration signal. Some examples may involve providing the one or more calibration signal parameters to each audio device of the plurality of audio devices. In some such implementations, determining the one or more calibration signal parameters may involve scheduling a time slot for each audio device of the plurality of audio devices to play a modified audio playback signal. In some examples, a first time slot for a first audio device may be different from a second time slot for a second audio device.
[0030] In some examples, determining the one or more calibration signal parameters may involve determining a frequency band for each audio device of the plurality of audio devices to reproduce the modified audio playback signal, and in some such examples, a first frequency band for a first audio device may be different from a second frequency band for a second audio device.
[0031] According to some examples, determining the one or more calibration signal parameters may involve determining a DSSS spreading code for each audio device of the plurality of audio devices. In some instances, a first spreading code for a first audio device may be different from a second spreading code for a second audio device. Some examples may involve determining at least one spreading code length based at least in part on the audibility of the corresponding audio device.
[0032] In some examples, determining the one or more calibration signal parameters may involve applying an acoustic model based at least in part on the interaudibility of each of a plurality of audio devices in the audio environment.
[0033] Some methods may involve determining that calibration signal parameters for an audio device are at a maximum robustness level. Some such methods may involve determining that a calibration signal from an audio device cannot be successfully extracted from a microphone signal. Some such methods may involve causing all other audio devices to mute at least a portion of the corresponding audio device playback. In some examples, this portion may be or include the calibration signal component.
[0034] Some implementations may involve simultaneously playing the modified audio playback signal on each of multiple audio devices in an audio environment.
[0035] According to some examples, at least a portion of the first audio playback signal, at least a portion of the second audio playback signal, or at least a portion of each of the first audio playback signal and the second audio playback signal corresponds to silence.
[0036] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. Thus, some innovative aspects of the subject matter described in this disclosure may be implemented via one or more non-transitory media having software stored thereon.
[0037] At least some aspects of the present disclosure may be implemented via an apparatus or system. For example, one or more devices may be capable of at least partially performing the methods disclosed herein. In some implementations, the apparatus is or includes an audio processing system having an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof.
[0038] The details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. It should be noted that the relative dimensions of the following figures may not be drawn to scale. [Brief explanation of the drawings]
[0039] Like reference numbers and designations in the various drawings indicate like elements.
[0040] [Figure 1A] An example of an audio environment is shown below.
[0041] [Figure 1B] FIG. 1 is a block diagram illustrating example components of an apparatus capable of implementing various aspects of the disclosure.
[0042] [Figure 2] FIG. 10 is a block diagram illustrating example audio device elements in accordance with some disclosed implementations.
[0043] [Figure 3] FIG. 10 is a block diagram illustrating an example of an audio device element in accordance with another disclosed implementation.
[0044] [Figure 4] FIG. 10 is a block diagram illustrating an example of an audio device element in accordance with another disclosed implementation.
[0045] [Figure 5] 1 is a graph showing an example of the levels of a content stream component reproduced by an audio device and a direct sequence spread spectrum (DSSS) signal component reproduced by an audio device over a range of frequencies.
[0046] [Figure 6] 10 is a graph showing an example of the power of two calibration signals having different bandwidths but located at the same center frequency.
[0047] [Figure 7] 1 illustrates elements of an example leadership module.
[0048] [Figure 8] 1 shows another example of an audio environment.
[0049] [Figure 9] FIG. 8 shows examples of acoustic calibration signals generated by audio devices 100B and 100C.
[0050] [Figure 10] 1 is a graph providing an example of a time domain multiple access (TDMA) method;
[0051] [Figure 11] 1 is a graph illustrating an example of a frequency domain multiple access (FDMA) method.
[0052] [Figure 12] 10 is a graph showing another example of a governance method.
[0053] [Figure 13]10 is a graph showing another example of a governance method.
[0054] [Figure 14] 1 illustrates elements of another example audio environment.
[0055] [Figure 15] FIG. 10 is a flow diagram outlining another example of the disclosed audio device control method.
[0056] [Figure 16] 1 shows another example of an audio environment.
[0057] [Figure 17] 1 is a block diagram illustrating an example of a calibration signal demodulator element, a baseband processor element, and a calibration signal generator element in accordance with some disclosed implementations.
[0058] [Figure 18] 10 illustrates elements of a calibration signal demodulator according to another example.
[0059] [Figure 19] FIG. 1 is a block diagram illustrating an example of a baseband processor element according to some disclosed implementations.
[0060] [Figure 20] 1 shows an example of a delayed waveform.
[0061] [Figure 21] 1 shows another example of an audio environment.
[0062] [Figure 22A] 1 is an example of a spectrogram of a modified audio playback signal.
[0063] [Figure 22B] 1 is a graph showing an example of a gap in the frequency domain.
[0064] [Figure 22C] 1 is a graph illustrating an example of a gap in the time domain.
[0065] [Figure 22D] 10 shows an example of a modified audio playback signal including coordinated gaps for multiple audio devices in an audio environment.
[0066] [Figure 23A] 10 is a graph showing an example of a filter response used to create a gap and a filter response used to measure the frequency domain of a microphone signal used during a measurement session.
[0067] [Figure 23B] 1 is a graph illustrating an example of a gap allocation strategy. [Figure 23C] 1 is a graph illustrating an example of a gap allocation strategy. [Figure 23D] 1 is a graph illustrating an example of a gap allocation strategy. [Figure 23E] 1 is a graph illustrating an example of a gap allocation strategy.
[0068] [Figure 24] 1 shows another example of an audio environment.
[0069] [Figure 25A] FIG. 1C is a flow diagram outlining an example of a method that may be performed by an apparatus such as that shown in FIG. 1B.
[0070] [Figure 25B] FIG. 2 is a block diagram of elements of an example embodiment configured to implement a zone classifier.
[0071] [Figure 26] 1 presents a block diagram of an example system for coordinated gap insertion.
[0072] [Figure 27A]1 illustrates the first half of a system block diagram showing example elements of a commanding device and a commanded audio device, according to some disclosed implementations. [Figure 27B] 10 shows the second half of a system block diagram illustrating example elements of a commanding device and a commanded audio device, according to some disclosed implementations.
[0073] [Figure 28] FIG. 10 is a flow diagram outlining another example of the disclosed audio device control method.
[0074] [Figure 29] FIG. 10 is a flow diagram outlining another example of the disclosed audio device control method.
[0075] [Figure 30] Examples of time-frequency assignments for calibration signals, gaps for noise estimation, and gaps for listening to a single audio device are shown.
[0076] [Figure 31] This example shows an audio environment, which is a living space.
[0077] [Figure 32] 1 is a block diagram illustrating three types of disclosed implementations. [Figure 33] 1 is a block diagram illustrating three types of disclosed implementations. [Figure 34] 1 is a block diagram illustrating three types of disclosed implementations.
[0078] [Figure 35] An example of a heat map is shown below.
[0079] [Figure 36] FIG. 10 is a block diagram illustrating an example of another implementation.
[0080] [Figure 37] FIG. 10 is a flow diagram outlining an example of another method that may be performed by an apparatus or system such as those disclosed herein.
[0081] [Figure 38] FIG. 1 is a block diagram illustrating an example of a system according to another implementation.
[0082] [Figure 39] FIG. 10 is a flow diagram outlining an example of another method that may be performed by an apparatus or system such as those disclosed herein.
[0083] [Figure 40] This example shows a floor plan for another audio environment in a living space.
[0084] [Figure 41] Shows an example of the geometric relationships between four audio devices in an environment.
[0085] [Figure 42] 42 shows an audio emitter positioned within the audio environment of FIG. 41.
[0086] [Figure 43] 42 shows an audio receiver positioned within the audio environment of FIG. 41.
[0087] [Figure 44] 1C is a flow diagram outlining another example of a method that may be performed by a control system of an apparatus such as that shown in FIG. 1B.
[0088] [Figure 45] 1 is a flow diagram outlining an example method for automatically estimating a device's position and orientation based on direction of arrival (DOA) data.
[0089] [Figure 46]1 is a flow diagram outlining an example method for automatically estimating a device's position and orientation based on DOA and time-of-arrival (TOA) data.
[0090] [Figure 47] FIG. 10 is a flow diagram outlining another example method for automatically estimating a device's position and orientation based on DOA and TOA data.
[0091] [Figure 48A] 1 shows another example of an audio environment.
[0092] [Figure 48B] 10 illustrates an example of determining listener angular orientation data.
[0093] [Figure 48C] 10 illustrates an additional example of determining listener angular orientation data.
[0094] [Figure 48D] FIG. 48C illustrates an example of determining the appropriate rotation of audio device coordinates according to the method described above.
[0095] [Figure 49] FIG. 10 is a flow diagram outlining another example of a localization method.
[0096] [Figure 50] FIG. 10 is a flow diagram outlining another example of a localization method.
[0097] [Figure 51] This example shows the floor plan of another listening environment, a living space.
[0098] [Figure 52] 10 is a graph of points illustrating speaker activation in an exemplary embodiment.
[0099] [Figure 53] 10 is a graph of trilinear interpolation between points showing speaker activations, according to an example.
[0100] [Figure 54] FIG. 10 is a block diagram of a minimal version of another embodiment.
[0101] [Figure 55] An alternative (more capable) embodiment with additional features is shown.
[0102] [Figure 56] FIG. 10 is a flow diagram outlining another example of the disclosed method. DETAILED DESCRIPTION OF THE INVENTION
[0103] To achieve a convincing spatial reproduction of media and entertainment content, the physical layout and relative capabilities of available speakers should be evaluated and considered. Similarly, to provide high-quality voice-driven interactions (with both virtual assistants and remote speakers), users need both to be heard and to hear the conversation being reproduced through speakers. As more collaborative devices are added to an audio environment, the combined usefulness to users is expected to increase, as it becomes more common for devices to be within convenient audio range. A greater number of speakers allows for a greater sense of immersion, as the spatiality of media presentation can be exploited.
[0104] Sufficient coordination and cooperation between devices could potentially allow these opportunities and experiences to be realized. Acoustic information about each audio device is a key component of such coordination and cooperation. Such acoustic information may include the audibility of each loudspeaker from various positions in the audio environment, as well as the amount of noise in the audio environment.
[0105] Some previous methods for mapping and calibrating a constellation of smart audio devices require a dedicated calibration procedure whereby known stimuli are played from the audio devices (often one audio device at a time) while one or more microphones record. While this process can be made to appeal to selected demographics of users through creative sound design, the need to repeatedly re-run the process as devices are added, removed, or simply rearranged presents a barrier to widespread adoption. Imposing such a procedure on users can be intrusive to the normal operation of the device and may frustrate some users.
[0106] An equally popular and more rudimentary approach is manual user intervention via a software application ("app") and / or a guided process in which the user indicates the physical location of audio devices within the audio environment. Such approaches present additional barriers to user adoption and may provide relatively less information to the system than a dedicated calibration procedure.
[0107] Calibration and mapping algorithms generally require some basic acoustic information about each audio device in the audio environment. Many such methods have been proposed, using a set of different basic acoustic measurements and measured acoustic properties. Examples of acoustic properties (also referred to herein as "acoustic scene metrics") derived from microphone signals for use in such algorithms include: · Estimates of the physical distance between devices (acoustic ranging); Estimate of the angle between devices (Direction of Arrival (DoA)); · Estimates of the impulse response between devices (e.g., through a swept sine wave stimulus or other measurement signal); · Estimated background noise.
[0108] However, existing calibration and mapping algorithms are generally not implemented to respond to changes in the acoustic scene of an audio environment, such as the movement of people within the audio environment or the repositioning of audio devices within the audio environment.
[0109] A coordinated system of smart audio devices as disclosed herein can provide users with the flexibility to place devices anywhere within their listening environment (also referred to herein as the audio environment). In some implementations, the audio devices are configured to automatically self-organize and calibrate.
[0110] Calibration may be conceptually divided into two or more layers. One such layer involves what may be referred to herein as “geometric mapping.” Geometric mapping may involve discovering the physical location and orientation of smart audio devices and one or more people within the audio environment. In some examples, geometric mapping may involve discovering the physical location of noise sources and / or legacy audio devices such as televisions (“TVs”) and sound bars. Geometric mapping is important for many reasons. For example, it is important that a flexible renderer be provided with accurate geometric mapping information in order to correctly render a sound scene. Conversely, legacy systems employing canonical loudspeaker layouts such as 5.1 have been designed under the assumption that the loudspeakers are placed in predetermined locations and that the listener is seated in a “sweet spot” facing the center loudspeaker and / or midway between the left and right front loudspeakers.
[0111] The second conceptual layer of calibration involves processing audio data (e.g., audio leveling and equalization) to account for loudspeaker manufacturing variations, room placement, and acoustic effects. In legacy cases, particularly soundbars and audio / video receivers (AVRs), users can optionally apply manual gain and EQ curves or connect dedicated reference microphones at the listening position for automatic calibration. However, it is known that a very small percentage of the population is willing to go this far. Therefore, a coordinated system of smart audio devices requires a method for automating audio processing (especially level and EQ calibration) without requiring the use of reference microphones at the listener position, a process sometimes referred to herein as "audibility mapping." Geometric mapping and audibility mapping are two major components of what may be referred to herein as "acoustic mapping."
[0112] This disclosure describes several techniques that can be used in various combinations to provide automated acoustic mapping. Acoustic mapping may be pervasive and ongoing. Such acoustic mapping is sometimes referred to as “continuous” in the sense that it may be continued after an initial setup process and may respond to changing conditions in the audio environment, such as changing noise sources and / or levels, loudspeaker relocation, deployment of additional loudspeakers, relocation and / or reorientation of one or more listeners, etc.
[0113] Some disclosed methods involve generating a calibration signal that is injected into (e.g., mixed with) audio content being rendered by an audio device in an audio environment. In some such examples, the calibration signal may be or include an acoustic direct sequence spread spectrum (DSSS) signal.
[0114] In other examples, the calibration signal may be or include other types of acoustic calibration signals, such as a swept sine wave acoustic signal, white noise, "colored noise" such as pink noise (a spectrum of frequencies whose intensity decreases at a rate of 3 decibels per octave), an acoustic signal corresponding to music, etc. Such a method may enable an audio device to generate observations after receiving calibration signals transmitted by other audio devices in the audio environment. In some implementations, each participating audio device in the audio environment may be configured to generate an acoustic calibration signal, inject the acoustic calibration signal into a rendered loudspeaker feed signal to generate a modified audio playback signal, and cause a loudspeaker system to play the modified audio playback signal to generate a first audio device playback sound. In some implementations, each participating audio device in the audio environment may be configured to simultaneously detect audio device playback sounds from other coordinated audio devices in the audio environment and process the audio device playback sounds to extract the acoustic calibration signal. Thus, although specific examples using acoustic DSSS signals are provided herein, these should be seen as specific examples within the broader category of acoustic calibration signals.
[0115] DSSS signals have previously been deployed in the context of telecommunications. When used in a telecommunications context, DSSS signals are used to spread transmitted data over a wider frequency range before the transmitted data is sent over a channel to a receiver. In contrast, most or all of the disclosed implementations do not involve using DSSS signals to modify or transmit data. Instead, such disclosed implementations involve sending DSSS signals between audio devices in an audio environment. What happens to the transmitted DSSS signal between the transmitter and receiver is itself the information being transmitted. This is one important difference between how DSSS signals are used in a telecommunications context and how DSSS signals are used in the disclosed implementations.
[0116] Additionally, disclosed implementations involve transmitting and receiving acoustic DSSS signals rather than transmitting and receiving electromagnetic DSSS signals. In many disclosed implementations, the acoustic DSSS signal is inserted into a content stream rendered for playback such that the acoustic DSSS signal is included in the played audio. According to some such implementations, the acoustic DSSS signal is inaudible to humans, so that a person in the audio environment does not perceive the acoustic DSSS signal but merely detects the audio content being played.
[0117] Another difference between the use of acoustic DSSS signals as disclosed herein and how DSSS signals are used in telecommunications contexts involves what may be referred to herein as the "near / far problem." In some instances, the acoustic DSSS signals disclosed herein may be transmitted and received by many audio devices in an audio environment. The acoustic DSSS signals may potentially overlap in time and frequency. Some disclosed implementations rely on how DSSS spreading codes are generated to separate the acoustic DSSS signals. In some instances, audio devices may be so close to each other that the signal levels may violate acoustic DSSS signal separation, making it difficult to separate the signals. This is one manifestation of the near-far problem, and several solutions for it are disclosed herein.
[0118] Some methods may include receiving a first content stream including a first audio signal, rendering the first audio signal to generate a first audio playback signal, generating a first calibration signal, generating a first modified audio playback signal by inserting the first calibration signal into the first audio playback signal, and playing the first modified audio playback signal on a loudspeaker system to generate a first audio device playback sound. The method may also include receiving microphone signals corresponding to at least the first audio device playback sound and second through Nth audio device playback sounds corresponding to second through Nth modified audio playback signals (including the second through Nth calibration signals) played by second through Nth audio devices, extracting the second through Nth calibration signals from the microphone signals, and estimating at least one acoustic scene metric based at least in part on the second through Nth calibration signals.
[0119] The acoustic scene metrics may be or may include audio device audibility, audio device impulse response, angle between audio devices, audio device position, and / or audio environment noise. Some disclosed methods may involve controlling one or more aspects of audio device playback based at least in part on the acoustic scene metrics.
[0120] Some disclosed methods may involve orchestrating multiple audio devices to perform a method involving a calibration signal. Some such methods may involve causing a first audio device in an audio environment to generate a first calibration signal, by a control system, causing the control system to insert the first calibration signal into a first audio playback signal corresponding to a first content stream to generate a first modified audio playback signal for the first audio device, and causing the control system to play the first modified audio playback signal to generate the first audio device playback sound.
[0121] Some such methods may include causing a second audio device in the audio environment to generate a second calibration signal by the control system, causing the control system to insert the second calibration signal into a second content stream to generate a second modified audio playback signal for the second audio device, and causing the control system to play the second modified audio playback signal to generate second audio device playback sound.
[0122] Some such implementations may involve causing at least one microphone in the audio environment to detect at least a first audio device and a second audio device and generate microphone signals corresponding to the at least a first audio device and a second audio device. Some such methods may involve causing the control system to extract at least a first calibration signal and a second calibration signal from the microphone signals, and causing the control system to estimate at least one acoustic scene metric based at least in part on the first calibration signal and the second calibration signal.
[0123] Figure 1A shows an example of an audio environment. As with other figures provided herein, the types and number of elements shown in Figure 1A are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements.
[0124] According to this example, audio environment 130 is a living space in a home. In the example shown in FIG. 1A , audio devices 100A, 100B, 100C, and 100D are located within audio environment 130. In this example, each of audio devices 100A-100D includes a corresponding one of loudspeaker systems 110A, 110B, 110C, and 110D. According to this example, loudspeaker system 110B of audio device 100B includes at least left loudspeaker 110B1 and right loudspeaker 110B2. In this case, audio devices 100A-100D include loudspeakers of different sizes and capabilities. At the time depicted in FIG. 1A , audio devices 100A-100D are generating corresponding instances of audio device playback sounds 120A, 120B1, 120B2, 120C, and 120D.
[0125] In this example, each of audio devices 100A-100D includes a corresponding one of microphone systems 111A, 111B, 111C, and 111D. Each of microphone systems 111A-111D includes one or more microphones. In some examples, audio environment 130 may include at least one audio device that lacks a loudspeaker system or at least one audio device that lacks a microphone system.
[0126] In some instances, at least one acoustic event may be occurring within audio environment 130. For example, one such acoustic event may be caused by a speaker, which in some instances may be issuing a voice command. In other instances, the acoustic event may be caused, at least in part, by a variable element, such as a door or window, in audio environment 130. For example, when a door opens, sounds from outside audio environment 130 may be perceived more clearly within audio environment 130. Furthermore, the changing angle of the door may change some of the echo paths within audio environment 130.
[0127] FIG. 1B is a block diagram illustrating example components of a device capable of implementing various aspects of the present disclosure. As with other figures provided herein, the types and number of elements shown in FIG. 1B are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements. According to some examples, device 150 may be configured to perform at least some of the methods disclosed herein. In some implementations, device 150 may be or include one or more components of an audio system. For example, device 150 may be an audio device, such as a smart audio device, in some implementations. In other examples, device 150 may be a mobile device (such as a cellular phone), a laptop computer, a tablet device, a television, or another type of device.
[0128] In the example shown in FIG. 1A, audio devices 100A-100D are instances of device 150. According to some examples, audio environment 100 of FIG. 1A may include an orchestrating device, such as what may be referred to herein as a smart home hub. The smart home hub (or other orchestrating device) may be an instance of device 150. In some implementations, one or more of audio devices 100A-100D may be capable of functioning as an orchestrating device.
[0129] According to some alternative implementations, apparatus 150 may be or include a server. In some such examples, apparatus 150 may be or include an encoder. Thus, in some instances, apparatus 150 may be a device configured for use in an audio environment, such as a home audio environment, while in other instances, apparatus 150 may be a device configured for use in the "cloud," e.g., a server.
[0130] In this example, device 150 includes interface system 155 and control system 160. Interface system 155, in some implementations, may include a wired or wireless interface configured to communicate with one or more other devices in the audio environment. The audio environment, in some examples, may be a home audio environment. In other examples, the audio environment may be another type of environment, such as an office environment, an automobile environment, a train environment, a street or sidewalk environment, a park environment, etc. Interface system 155, in some implementations, may be configured to exchange control information and associated data with audio devices in the audio environment. The control information and associated data, in some examples, may relate to one or more software applications running on device 150.
[0131] In some implementations, interface system 155 may be configured to receive or provide a content stream. The content stream may include audio data. The audio data may include, but is not limited to, an audio signal. In some instances, the audio data may include spatial data, such as channel data and / or spatial metadata. The metadata may be provided, for example, by what may be referred to herein as an “encoder.” In some examples, the content stream may include video data and audio data corresponding to the video data.
[0132] Interface system 155 may include one or more network interfaces and / or one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). According to some implementations, interface system 155 may include one or more wireless interfaces configured for, for example, Wi-Fi or Bluetooth® communications.
[0133] Interface system 155, in some examples, may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, interface system 155 may include one or more interfaces between control system 160 and a memory system, such as optional memory system 165 shown in FIG. 1B . However, control system 160 may include a memory system in some cases. Interface system 155, in some implementations, may be configured to receive input from one or more microphones in the environment.
[0134] In some implementations, control system 160 may be configured to at least partially perform the methods disclosed herein and may include, for example, a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.
[0135] In some implementations, control system 160 may reside in more than one device. For example, in some implementations, a portion of control system 160 may reside in a device within one of the environments depicted herein, and another portion of control system 160 may reside in a device outside the environment, such as a server, a mobile device (e.g., a smartphone or tablet computer), or the like. In other examples, a portion of control system 160 may reside in a device within one of the environments depicted herein, and another portion of control system 160 may reside in one or more other devices in the environment. For example, control system functionality may be distributed across multiple smart audio devices in the environment, or may be shared by a coordinating device (such as what may be referred to herein as a smart home hub) and one or more other devices in the environment. In other examples, a portion of control system 160 may reside in a device implementing a cloud-based service, such as a server, and another portion of control system 160 may reside in another device implementing the cloud-based service, such as another server, memory device, or the like. Interface system 155 may also reside in more than one device, in some examples.
[0136] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may reside, for example, in optional memory system 165 and / or control system 160 shown in FIG. 1B. Thus, various innovative aspects of the subject matter described in this disclosure may be implemented in one or more non-transitory media having software stored thereon. The software may include, for example, instructions for controlling at least one device to perform some or all of the methods disclosed herein. The software may be executable by one or more components of a control system, such as, for example, control system 160 of FIG. 1B.
[0137] In some examples, device 150 may include optional microphone system 111 shown in FIG. 1B . Optional microphone system 111 may include one or more microphones. According to some examples, optional microphone system 111 may include an array of microphones. The microphone array may, in some cases, be configured for receive-side beamforming, e.g., according to instructions from control system 160. In some examples, the microphone array may be configured to determine direction of arrival (DOA) and / or time of arrival (TOA) information, e.g., according to instructions from control system 160. Alternatively or additionally, control system 160 may be configured to determine direction of arrival (DOA) and / or time of arrival (TOA) information, e.g., according to microphone signals received from microphone system 111.
[0138] In some implementations, one or more of the microphones may be part of or associated with another device, such as a speaker of a speaker system, a smart audio device, etc. In some examples, device 150 may not include microphone system 111. However, in some such implementations, device 150 may still be configured to receive microphone data for one or more microphones in the audio environment via interface system 160. In some such implementations, a cloud-based implementation of device 150 may be configured to receive microphone data or data corresponding to microphone data from one or more microphones in the audio environment via interface system 160.
[0139] According to some implementations, device 150 may include optional loudspeaker system 110 shown in FIG. 1B. Optional loudspeaker system 110 may include one or more loudspeakers, sometimes referred to herein as "speakers" or more generally as "audio reproduction transducers." In some examples (e.g., cloud-based implementations), device 150 may not include loudspeaker system 110.
[0140] In some implementations, device 150 may include optional sensor system 180 shown in FIG. 1B . Optional sensor system 180 may include one or more touch sensors, gesture sensors, motion detectors, etc. According to some implementations, optional sensor system 180 may include one or more cameras. In some implementations, the camera(s) may be freestanding cameras. In some examples, one or more cameras of optional sensor system 180 may reside within a smart audio device, which may be a single-purpose audio device or a virtual assistant. In some such examples, one or more cameras of optional sensor system 180 may reside within a television, a mobile phone, or a smart speaker. In some examples, device 150 may not include sensor system 180. However, in some such implementations, device 150 may still be configured to receive sensor data for one or more sensors in the audio environment via interface system 160.
[0141] In some implementations, device 150 may include optional display system 185 shown in FIG. 1B . Optional display system 185 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some cases, optional display system 185 may include one or more organic light-emitting diode (OLED) displays. In some examples, optional display system 185 may include one or more displays of a smart audio device. In other examples, optional display system 185 may include a television display, a laptop display, a mobile device display, or another type of display. In some examples in which device 150 includes display system 185, sensor system 180 may include a touch sensor system and / or a gesture sensor system proximate one or more displays of display system 185. According to some such implementations, control system 160 may be configured to control display system 185 to present one or more graphical user interfaces (GUIs).
[0142] According to some such examples, device 150 may be or include a smart audio device. In some such implementations, device 150 may be or include a wake word detector. For example, device 150 may be or include a virtual assistant.
[0143] FIG. 2 is a block diagram illustrating example audio device elements according to some disclosed implementations. As with other figures provided herein, the types and number of elements shown in FIG. 2 are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements. In this example, audio device 100A of FIG. 2 is an instance of apparatus 150 described above with reference to FIG. 1B. In this example, audio device 100A is one of multiple audio devices in an audio environment and, in some instances, may be an example of audio device 100A shown in FIG. 1A. In this example, the audio environment includes at least two other dominant audio devices, audio device 100B and audio device 100C.
[0144] According to this implementation, audio device 100A includes the following elements: 110A: an instance of the loudspeaker system 110 of FIG. 1B including one or more loudspeakers; 111A: an instance of microphone system 111 of FIG. 1B including one or more microphones; 120A, B, C: Audio device playback sounds corresponding to rendered content being played by audio devices 100A-100C in the same acoustic space; 201A: Audio playback signal output by the rendering module 210A; 202A: modified audio playback signal output by calibration signal injector 211A; 203A: calibration signal output by calibration signal generator 212A; 204A: Calibration signal replica corresponding to calibration signals generated by other audio devices in the audio environment (in this example, at least audio devices 100B and 100C). In some examples, calibration signal replica 204A may be received (e.g., via a wireless communication protocol such as Wi-Fi or Bluetooth®) from an external source, such as a commanding device (which may be another audio device in the audio environment, another local device such as a smart home hub, etc.); 205A: Calibration information associated with and / or used by one or more audio devices in an audio environment. Calibration information 205A may include parameters used by control system 160 of audio device 100A to generate a calibration signal, modulate a calibration signal, demodulate a calibration signal, etc. Calibration information 205A may include one or more DSSS spreading code parameters and one or more DSSS carrier parameters, in some examples. DSSS spreading code parameters may include, for example, DSSS spreading code length information, chipping rate information (or chip period information), etc. One chip period is the time it takes for one chip (bit) of the spreading code to be played. The reciprocal of the chip period is the chipping rate. Bits in a DSSS spreading code may be referred to as "chips" to indicate that they do not contain data (as bits normally contain). In some instances, DSSS spreading code parameters may include a pseudorandom number sequence. The calibration information 205A may, in some examples, indicate which audio device is generating the acoustic calibration signal. In some examples, the calibration information 205A may be received (e.g., via wireless communication) from an external source, such as a master device; 206A: microphone signal received by microphone 111A; 208A: demodulated coherent baseband signal; 210A: a rendering module configured to render an audio signal of a content stream, such as audio data for music, movies, and television programs, to generate an audio playback signal; 211A: a calibration signal injector configured to insert the calibration signal 230A modulated by the calibration signal modulator 220A into the audio playback signal generated by the rendering module 210A to generate a modified audio playback signal. The insertion process may, for example, be a mixing process in which the calibration signal 230A modulated by the calibration signal modulator 220A is mixed with the audio playback signal generated by the rendering module 210A to generate the modified audio playback signal; 212A: A calibration signal generator configured to generate a calibration signal 203A and provide the calibration signal 203A to the calibration signal modulator 220A and the calibration signal demodulator 214A. In some examples, the calibration signal generator 212A may include a DSSS spreading code generator and a DSSS carrier generator. In this example, the calibration signal generator 212A provides a calibration signal replica 204A to the calibration signal demodulator 214A; 214A: An optional calibration signal demodulator configured to demodulate the microphone signal 206A received by the microphone 111A. In this example, the calibration signal demodulator 214A outputs a demodulated coherent baseband signal 208A. Demodulation of the microphone signal 206A can be performed using standard correlation techniques, including, for example, an integrate and dump style matched filtering correlator bank. Some detailed examples are provided below. To improve the performance of these demodulation techniques, in some implementations, the microphone signal 206A can be filtered before demodulation to remove unwanted content / phenomena. According to some implementations, the demodulated coherent baseband signal 208A can be filtered before being provided to the baseband processor 218A. The signal-to-noise ratio (SNR) generally improves as the integration time increases (as the length of the spreading code used increases). Not all types of calibration signals (e.g., white noise and acoustic signals corresponding to music) require modulation before being mixed with the rendered audio data for playback, so some implementations may not include a calibration signal demodulator; 218A: A baseband processor configured for baseband processing of the demodulated coherent baseband signal 208A. In some examples, the baseband processor 218A may be configured to implement techniques such as incoherent averaging to improve the SNR by reducing the variance of the squared waveform that generates the delayed waveform. Some detailed examples are provided below. In this example, the baseband processor 218A is configured to output one or more estimated acoustic scene metrics 225A; 220A: An optional calibration signal modulator configured to modulate the calibration signal 203A generated by the calibration signal generator to generate calibration signal 230A. As noted elsewhere herein, not all types of calibration signals require modulation before being mixed with rendered audio data for playback. Thus, some implementations may not include a calibration signal modulator; 225A: One or more observations derived from the calibration signal(s), also referred to herein as acoustic scene metrics. The acoustic scene metrics 225A may include or be data corresponding to time of flight, time of arrival, range, audio device audibility, audio device impulse response, angle between audio devices, audio device position, audio environment noise, and / or signal-to-noise ratio; 233A: Acoustic scene metric processing module, configured to receive and apply acoustic scene metrics 225A. In this example, acoustic scene metric processing module 233A is configured to generate information 235A (and / or commands) based at least in part on at least one acoustic scene metric 225A and / or at least one audio device characteristic. The audio device characteristic may correspond to audio device 100A or another audio device in the audio environment, depending on the particular implementation. The audio device characteristic may be stored in a memory of control system 160 or accessible to control system 210, for example; 235A: Information for controlling one or more aspects of audio processing and / or audio device playback. Information 235A may include, for example, information (and / or commands) for controlling a rendering process, an audio environment mapping process (such as an audio device auto-location process), an audio device calibration process, a noise suppression process, and / or an echo attenuation process.
[0145] Acoustic Scene Metrics Examples As mentioned above, in some implementations, the baseband processor 218A (or another module of the control system 160) may be configured to determine one or more acoustic scene metrics 225A. The following are some examples of acoustic scene metrics 225A:
[0146] Ranging A calibration signal received by an audio device from another audio device contains information about the distance between the two devices in the form of the signal's time of flight (ToF). According to some examples, the control system may be configured to extract delay information from the demodulated calibration signal and convert the delay information into a pseudorange measurement, for example, as follows: ρ=τc
[0147] In the above equation, τ represents delay information (also referred to herein as Time of Flight), ρ represents a pseudo-range measurement, and c represents the speed of sound. The term "pseudo-range" is used because range itself is not measured directly; therefore, range between devices is estimated according to timing estimates. In a distributed, asynchronous system of audio devices, each audio device operates on its own clock, and therefore biases exist in the raw delay measurements. Given a sufficient set of delay measurements, it is possible to resolve these biases, and sometimes even estimate them. Detailed examples of extracting delay information, generating and using pseudo-range measurements, and determining and resolving clock biases are provided below.
[0148] DoA Similar to ranging, using multiple microphones available on the listening device, the control system can be configured to estimate the direction of arrival (DoA) by processing the demodulated acoustic calibration signal. In some such implementations, the resulting DoA information can be used as input to a DoA-based audio device automatic localization method.
[0149] Audibility The signal strength of the demodulated acoustic calibration signal is proportional to the audibility of the audio device being listened to in the band over which the audio device is transmitting the acoustic calibration signal. In some implementations, the control system may be configured to make multiple observations across a range of frequency bands to obtain a banded estimate of the entire frequency range. With knowledge of the digital signal level of the transmitting audio device, the control system may, in some examples, be configured to estimate the absolute acoustic gain of the transmitting audio device.
[0150] FIG. 3 is a block diagram illustrating example audio device elements according to another disclosed implementation. As with other figures provided herein, the types and number of elements shown in FIG. 3 are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements. In this example, audio device 100A of FIG. 3 is an instance of apparatus 150 described above with reference to FIGS. 1B and 2. However, according to this implementation, audio device 100A is configured to coordinate multiple audio devices in an audio environment, including at least audio devices 100B, 100C, and 100D.
[0151] The implementation shown in Figure 3 includes all of the elements of Figure 2, as well as some additional elements. Elements common to Figures 2 and 3 will not be described again here, except to the extent their functionality may differ in the implementation of Figure 3. According to this implementation, audio device 100A includes the following elements and functionality: 120A, B, C, D: Audio device playback sounds corresponding to the rendered content being played by audio devices 100A-100D in the same acoustic space; 204A, B, C, D: calibration signal replicas corresponding to calibration signals generated by other audio devices in the audio environment (in this example, at least audio devices 100B, 100C, and 100D). In this example, calibration signal replicas 204A-204D are provided by governing module 213A, which provides calibration information 204B-204D to audio devices 100B-100D, for example, via wireless communication; 205A, B, C, D: These elements correspond to calibration information associated with and / or used by each of audio devices 100A-100D. Calibration information 205A may include parameters (e.g., one or more DSSS spreading code parameters and one or more DSSS carrier parameters) used by control system 160 of audio device 100A to generate a calibration signal, modulate a calibration signal, demodulate a calibration signal, etc. Calibration information 205B, 205C, and 205D may include parameters (e.g., one or more DSSS spreading code parameters and one or more DSSS carrier parameters) used by audio devices 100B, 100C, and 100D, respectively, to generate a calibration signal, modulate a calibration signal, demodulate a calibration signal, etc. Calibration information 205A-205D may, in some examples, indicate which audio device is generating the acoustic calibration signal; 213A: Coordination module. In this example, the coordination module 213A generates calibration information 205A-205D, provides calibration information 205A to the calibration signal generator 212A, provides calibration information 205A-205D to the calibration signal demodulator, and provides calibration information 205B-205D to the audio devices 100B-100D, e.g., via wireless communication. In some examples, the coordination module 213A generates calibration information 205A-205D based at least in part on information 235A-235D and / or acoustic scene metrics 225A-225D; 214A: A calibration signal demodulator configured to demodulate microphone signal 206A received by at least microphone 111A. In this example, calibration signal demodulator 214A outputs demodulated coherent baseband signal 208A. In some alternative implementations, calibration signal demodulator 214A may receive and demodulate microphone signals 206B-206D from audio devices 100B-100D and output demodulated coherent baseband signals 208B-208D; 218A: A baseband processor configured for baseband processing of at least the demodulated coherent baseband signal 208A, and in some examples, the demodulated coherent baseband signals 208B-208D received from the audio devices 100B-100D. In this example, the baseband processor 218A is configured to output one or more estimated acoustic scene metrics 225A-225D. In some implementations, the baseband processor 218A is configured to determine the acoustic scene metrics 225B-225D based on the demodulated coherent baseband signals 208B-208D received from the audio devices 100B-100D. However, in some cases, the baseband processor 218A (or the acoustic scene metric processing module 233A) may receive the acoustic scene metrics 225B-225D from the audio devices 100B-100D; 233A: Acoustic scene metric processing module, configured to receive and apply acoustic scene metrics 225A-225D. In this example, acoustic scene metric processing module 233A is configured to generate information 235A-235D based at least in part on acoustic scene metrics 225A-225D and / or at least one audio device characteristic, which may correspond to one or more of audio device 100A and / or audio devices 100B-100D.
[0152] FIG. 4 is a block diagram illustrating example audio device elements according to another disclosed implementation. As with other figures provided herein, the types and number of elements shown in FIG. 4 are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements. In this example, audio device 100A of FIG. 4 is an instance of apparatus 150 described above with reference to FIGS. 1B, 2, and 3. The implementation shown in FIG. 4 includes all of the elements of FIG. 3, as well as additional elements. Elements common to FIGS. 2 and 3 will not be described again here, except to the extent their functionality may differ in the implementation of FIG. 4.
[0153] According to this implementation, control system 160 is configured to process received microphone signal 206A to generate preprocessed microphone signal 207A. In some implementations, processing the received microphone signal may involve applying a bandpass filter and / or echo cancellation. In this example, control system 160 (more specifically, calibration signal demodulator 214A) is configured to extract a calibration signal from preprocessed microphone signal 207A.
[0154] According to this example, microphone system 111A includes an array of microphones, which in some cases may be or include one or more directional microphones. In this implementation, processing the received microphone signals involves receive-side beamforming, in this example via beamformer 215A. In this example, preprocessed microphone signal 207A output by beamformer 215A is or includes a spatial microphone signal.
[0155] In this implementation, calibration signal demodulator 214A processes spatial microphone signals, which can improve performance for audio systems in which audio devices are spatially distributed around the audio environment. Receiver beamforming is one way to circumvent the aforementioned "near-far problem"; for example, control system 160 can be configured to use beamforming to receive audio device playback from more distant and / or quieter audio devices, compensating for closer and / or louder audio devices.
[0156] Receiver beamforming may involve, for example, delaying and multiplying the signal from each microphone in the microphone array by a different factor. In some examples, the beamformer 215A may apply a Dolph-Chebyshev weighting pattern. However, in other implementations, the beamformer 215A may apply a different weighting pattern. According to some such examples, a main lobe may be generated along with nulls and side lobes. In addition to controlling the main lobe width (beamwidth) and side lobe levels, in some examples, the location of the nulls may be controlled.
[0157] Sub-audible signals According to some implementations, the calibration signal components of the audio device playback may not be audible to people in the audio environment. In some such implementations, the content stream components of the audio device playback may cause perceptual masking of the calibration signal components of the audio device playback.
[0158] 5 is a graph illustrating an example of the levels of a content stream component reproduced by an audio device and a DSSS signal component reproduced by an audio device over a frequency range, where curve 501 corresponds to the level of the content stream component and curve 530 corresponds to the level of the DSSS signal component.
[0159] A DSSS signal typically includes data, a carrier signal, and a spreading code. If we omit the need to transmit data over a channel, we can express the modulated signal s(t) as:
[0160] s(t)=AC(t)sin(2πf0t) where A represents the amplitude of the DSSS signal, C(t) represents the spreading code, and Sin() represents a sinusoidal carrier with a carrier frequency of f Hz. Curve 530 in Figure 5 corresponds to an example of s(t) in the above equation.
[0161] One potential advantage of some disclosed implementations that include acoustic DSSS signals is that spreading the signal can reduce the perceptibility of DSSS signal components in audio device playback, since the amplitude of the DSSS signal components is reduced for a given amount of energy in the acoustic DSSS signal.
[0162] This allows the DSSS signal component of the audio device playback sound (e.g., as represented by curve 530 in FIG. 5) to be at a level sufficiently lower than the level of the content stream component of the audio device playback sound (e.g., as represented by curve 501 in FIG. 5) so that the DSSS signal component is not perceptible to the listener.
[0163] Some disclosed implementations exploit the masking properties of the human auditory system to optimize parameters of the calibration signal in a manner that maximizes the signal-to-noise ratio (SNR) of derived calibration signal observations and / or reduces the probability of perception of the calibration signal components. Some disclosed examples involve applying weights to the levels of content stream components and / or applying weights to the levels of calibration signal components. Some such examples apply noise compensation methods, where the acoustic calibration signal components are treated as signals and the content stream components are treated as noise. Some such examples involve applying one or more weights according to (e.g., proportionally) a playback / listening objective metric.
[0164] DSSS spreading code As noted elsewhere herein, in some examples, the calibration information 205 provided by the governing device (e.g., provided by the governing module 213A described above with reference to FIG. 3) may include one or more DSSS spreading code parameters.
[0165] The spreading codes used to spread the carrier to generate the DSSS signal can be important. The set of DSSS spreading codes is preferably selected so that the corresponding DSSS signal has the following properties: 1. Sharp main lobe in the autocorrelation waveform; 2. Low sidelobes at non-zero delays in the autocorrelation waveform; 3. Low cross-correlation between any two spreading codes in said set of spreading codes used when multiple devices access the medium simultaneously (e.g., to simultaneously play back modified audio playback signals that include DSSS signal components); 4. The DSSS signal is unbiased (has a DC component of 0).
[0166] A family of spreading codes (e.g., Gold codes commonly used in GPS contexts) typically characterizes the above four points. When multiple audio devices are all simultaneously playing a modified audio playback signal that includes a DSSS signal component, and each audio device uses a different spreading code (one with good cross-correlation properties, e.g., low cross-correlation), the receiving audio device should be able to simultaneously receive and process all of the acoustic DSSS signals by using a code-domain multiple access (CDMA) method. By using a CDMA method, multiple audio devices can simultaneously transmit acoustic DSSS signals, possibly using a single frequency band. The spreading codes may be generated during runtime and / or pre-generated and stored in memory, e.g., in a data structure such as a look-up table.
[0167] To implement DSSS, in some examples, binary phase shift keying (BPSK) modulation may be utilized. Furthermore, DSSS spreading codes may in some examples be made orthogonal (interplexed) with one another to implement a quadrature phase shift keying (QPSK) system, for example, as follows: s(t)=A I C I (t)cos(2πf0t)+A Q C Q (t)sin(2πf0t)
[0168] In the above formula, A I and A Q represent the amplitudes of the in-phase and quadrature signals, respectively, and C I and C Qwhere f represents the code sequences for the in-phase and quadrature signals, respectively, and f represents the center frequency (8200) of the DSSS signal. The above are example coefficients for parameterizing the DSSS carrier and DSSS spreading code, according to some examples. These parameters are examples of the calibration signal information 205 described above. As mentioned above, the calibration signal information 205 may be provided by a governing device, such as governing module 213A, and may be used by signal generator block 212, for example, to generate the DSSS signal.
[0169] 6 is a graph illustrating example powers of two calibration signals having different bandwidths but located at the same center frequency. In these examples, FIG. 6 shows the spectra of two calibration signals 630A and 630B, both centered at the same center frequency 605. In some examples, calibration signal 630A may be generated by one audio device in the audio environment (e.g., by audio device 100A), and calibration signal 630B may be generated by another audio device in the audio environment (e.g., by audio device 100B).
[0170] According to this example, calibration signal 630B is chipped at a higher rate than calibration signal 630A (i.e., more bits per second are used in the spread signal), resulting in a larger bandwidth 610B for calibration signal 630B than the bandwidth 610A for calibration signal 630A. For a given amount of energy for each calibration signal, the larger the bandwidth of calibration signal 630B, the lower the amplitude and perceptibility of calibration signal 630B relative to calibration signal 630A. A higher-bandwidth calibration signal also results in higher delay-resolution of the baseband data products, leading to higher-resolution estimates of acoustic scene metrics based on the calibration signal (such as time-of-flight estimates, time-of-arrival (ToA) estimates, range estimates, and direction-of-arrival (DoA) estimates). However, a higher-bandwidth calibration signal also increases the noise bandwidth of the receiver, thereby reducing the SNR of the extracted acoustic scene metrics. Furthermore, if the bandwidth of the calibration signal is too large, there may be coherence and fading problems associated with the calibration signal.
[0171] The length of the spreading codes used to generate the DSSS signals limits the amount of cross-correlation cancellation. For example, a 10-bit Gold code has only -26 dB rejection of adjacent codes. This can create an instance of the near-far problem mentioned above, where a relatively low-amplitude signal can be obscured by the cross-correlated noise of another, louder signal. Similar problems can arise with other types of calibration signals. Some of the novelty of the systems and methods described in this disclosure involves discipline schemes designed to mitigate or avoid such problems.
[0172] Orchestration Methods FIG. 7 illustrates elements of a governance module according to one example. As with other figures provided herein, the types and number of elements shown in FIG. 7 are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements. According to some examples, governance module 213 may be implemented by an instance of device 150 described above with reference to FIG. 1B. In some such examples, governance module 213 may be implemented by an instance of control system 160. In some examples, governance module 213 may be an instance of the governance module described above with reference to FIG. 3.
[0173] According to this implementation, the governance module 213 includes a perceptual model application module 710 , an acoustic model application module 711 , and an optimization module 712 .
[0174] In this example, the perceptual model application module 710 is configured to apply a model of the human auditory system to obtain one or more perceptual impact estimates 702 of the perceptual impact of the acoustic calibration signal on a listener in an acoustic space based at least in part on the a priori information 701. The acoustic space may be, for example, an audio environment in which an audio device governed by the governing module 213 is located, a room in such an audio environment, etc. The estimate(s) 702 may change over time. The perceptual impact estimates 702 may, in some examples, be estimates of a listener's ability to perceive the acoustic calibration signal based, for example, on the type and level of audio content (if any) currently being played in the acoustic space. The perceptual model application module 710 may be configured to apply one or more models of auditory masking, such as, for example, masking as a function of frequency and loudness, spatial auditory masking, etc. The perceptual model application module 710 may be configured to apply, for example, one or more models of human loudness perception, for example, human loudness perception as a function of frequency.
[0175] According to some examples, the a priori information 701 may be or include information related to the acoustic space, information related to the transmission of acoustic calibration signals in the acoustic space, and / or information related to listeners known to use the acoustic space. For example, the a priori information 701 may include information regarding the number of audio devices (e.g., of governed audio devices) in the acoustic space, the locations of the audio devices, the loudspeaker system and / or microphone system capabilities of the audio devices, information regarding the impulse response of the audio environment, information regarding one or more doors and / or windows of the audio environment, information regarding audio content currently being played in the acoustic space, etc. In some instances, the a priori information 701 may include information regarding the hearing ability of one or more listeners.
[0176] In this implementation, the acoustic model application module 711 is configured to obtain one or more acoustic calibration signal performance estimates 703 for acoustic calibration signals in the acoustic space based at least in part on the a priori information 701. For example, the acoustic model application module 711 may be configured to estimate how well each microphone system of an audio device can detect acoustic calibration signals from other audio devices in the acoustic space, which may be referred to herein as an aspect of the “mutual audibility” of the audio devices. Such mutual audibility may, in some cases, have been an acoustic scene metric previously estimated by the baseband processor based at least in part on previously received acoustic calibration signals. In some such implementations, the mutual audibility estimates may be part of the a priori information 701, and in some such implementations, the governing module 213 may not include the acoustic model application module 711. However, in some implementations, the mutual audibility estimation may be performed independently by the acoustic model application module 711.
[0177] In this example, the optimization module 712 is configured to determine calibration parameters 705 for all audio devices governed by the governance module 213 based at least in part on the perceptual impact estimates 702 and acoustic calibration signal performance estimates 703, and current play / listen objective information 704. The current play / listen objective information 704 may indicate, for example, the relative need for new acoustic scene metrics based on the acoustic calibration signal.
[0178] For example, if one or more audio devices are newly powered up within the acoustic space, there may be a high level of need for new acoustic scene metrics related to audio device auto-location, audio device inter-audibility, etc. At least some of the new acoustic scene metrics may be based on acoustic calibration signals. Similarly, if existing audio devices are moved within the acoustic space, there may be a high level of need for new acoustic scene metrics. Similarly, if a new noise source is located in or near the acoustic space, there may be a high level of need for determining new acoustic scene metrics.
[0179] If the current playback / listening intent information 704 indicates a high level of need for determining new acoustic scene metrics, the optimization module 712 may be configured to determine the calibration parameters 705 by placing a relatively higher weight on the acoustic calibration signal performance estimate 703 than on the perceptual impact estimate 702. For example, the optimization module 712 may be configured to determine the calibration parameters 705 by emphasizing the system's ability to generate high SNR observations of the acoustic calibration signal and de-emphasizing the impact / perceptibility of the acoustic calibration signal by the user. In some such examples, the calibration parameters 705 may correspond to an audible acoustic calibration signal.
[0180] However, if there have been no recent changes detected in or near the acoustic space and there are at least initial estimates of one or more acoustic scene metrics, there may not be a high level of need for new acoustic scene metrics. If there have been no recent changes detected in or near the acoustic space and there are at least initial estimates of one or more acoustic scene metrics, and audio content is currently playing within the acoustic space, the relative importance of immediately estimating one or more new acoustic scene metrics may be further reduced.
[0181] If the current playback / listening intent information 704 indicates a low-level need to determine new acoustic scene metrics, the optimization module 712 may be configured to determine the calibration parameters 705 by placing a relatively lower weight on the acoustic calibration signal performance estimate 703 than on the perceptual impact estimate 702. In such examples, the optimization module 712 may be configured to determine the calibration parameters 705 by de-emphasizing the system's ability to generate high SNR observations of the acoustic calibration signal and by emphasizing the impact / perceptibility of the acoustic calibration signal by a user. In some such examples, the calibration parameters 705 may correspond to a sub-audible acoustic calibration signal.
[0182] As discussed later in this document (e.g., in other examples of audio device control), the parameters of the acoustic calibration signal provide a wealth of variety in how the control device can modify the acoustic calibration signal to improve the performance of the audio system.
[0183] FIG. 8 shows another example of an audio environment. In FIG. 8, audio devices 100B and 100C are separated from device 100A by distances 810 and 811, respectively. In this particular situation, distance 811 is greater than distance 810. Assuming that audio devices 100B and 100C are producing audio device playback sounds at approximately the same level, this means that audio device 100A will receive the acoustic calibration signal from audio device 100C at a lower level than the acoustic calibration signal from audio device 100B due to the additional acoustic loss caused by the longer distance 811. In some embodiments, audio devices 100B and 100C may be coordinated to improve audio device 100A's ability to extract the acoustic calibration signal and determine acoustic scene metrics based on the acoustic calibration signal.
[0184] FIG. 9 shows an example of acoustic calibration signals generated by audio devices 100B and 100C of FIG. 8. In this example, these acoustic calibration signals have the same bandwidth and are located at the same frequency, but have different amplitudes. Here, acoustic calibration signal 230B is generated by audio device 100B, and the main lobe of acoustic calibration signal 230C is generated by audio device 100C. According to this example, the peak power of acoustic calibration signal 230B is 905B, and the peak power of acoustic calibration signal 230C is 905C. Here, acoustic calibration signal 230B and acoustic calibration signal 230C have the same center frequency 901.
[0185] In this example, the governing device (which in some examples may include an instance of governing module 213 of FIG. 7 and, in some instances, may be audio device 100A of FIG. 8 ) enhances the ability of audio device 100A to extract the acoustic calibration signal by equalizing the digital levels of the acoustic calibration signals generated by audio devices 100B and 100C. Here, the equalization causes the peak power of acoustic calibration signal 230C to be greater than the peak power of acoustic calibration signal 230B by a factor that offsets the difference in acoustic loss due to the difference between distances 810 and 811. Thus, according to this example, audio device 100A receives acoustic calibration signal 230B from audio device 100C at approximately the same level as the acoustic calibration signal received from audio device 100B due to the additional acoustic loss caused by the longer distance 811.
[0186] The area of the surface around a point sound source increases with the square of the distance from the source. This means that the same sound energy from the source is distributed over a larger area, and the energy intensity decreases with the square of the distance from the source, according to the inverse square law. If distance 810 is b and distance 811 is c, then the sound energy received by audio device 100A from audio device 100B is 1 / b 2 The sound energy received by audio device 100A from audio device 100C is proportional to 1 / c 2 The difference in sound energy is proportional to 1 / (c 2 -b 2 ) is proportional to the energy generated by the audio device 100C. 2 -b 2 ) times. This is an example of how the calibration parameters can be changed to improve performance.
[0187] In some implementations, the optimization process may be more complex and may take into account more factors than the inverse square law. In some examples, equalization may be performed via a full-band gain applied to the calibration signal or via an equalization (EQ) curve that allows for equalization of a non-flat (frequency-dependent) response of microphone system 110A.
[0188] FIG. 10 is a graph providing an example of a time-domain multiple access (TDMA) method. One way to avoid the near-far problem is to coordinate multiple audio devices transmitting and receiving acoustic calibration signals so that each audio device is scheduled for a different time slot to play its acoustic calibration signal. This is known as a TDMA method. In the example shown in FIG. 10, the coordinating device causes audio devices 1, 2, and 3 to emit acoustic calibration signals according to the TDMA method. In this example, audio devices 1, 2, and 3 emit acoustic calibration signals in the same frequency band. According to this example, the coordinating device causes audio device 3 to emit an acoustic calibration signal from time t0 to time t1, then the coordinating device causes audio device 2 to emit an acoustic calibration signal from time t1 to time t2, then the coordinating device causes audio device 1 to emit an acoustic calibration signal from time t2 to time t3, and so on.
[0189] Thus, in this example, no two calibration signals are transmitted or received simultaneously. Therefore, the remaining calibration signal parameters, such as amplitude, bandwidth, and length, are not relevant for multiple access (as long as each calibration signal remains within its assigned time slot). However, such calibration signal parameters are still relevant to the quality of the observations extracted from the calibration signals.
[0190] FIG. 11 is a graph illustrating an example of a frequency-domain multiple access (FDMA) method. In some implementations (e.g., due to limited bandwidth of the calibration signal), a command device may be configured to cause an audio device to simultaneously receive acoustic calibration signals from two other audio devices within the audio environment. In some such examples, if each audio device transmitting an acoustic calibration signal reproduces its respective acoustic calibration signal in a different frequency band, the acoustic calibration signals will differ significantly in received power level. This is an FDMA method. In the example FDMA method shown in FIG. 11, calibration signals 230B and 230C are transmitted simultaneously by different audio devices but have different center frequencies (f1 and f2) and are in different frequency bands (b1 and b2). In this example, the main lobe frequency bands b1 and b2 do not overlap. Such an FDMA method may be advantageous for situations in which the acoustic calibration signals have large differences in acoustic losses associated with their paths.
[0191] In some implementations, the command device may be configured to modify the FDMA, TDMA, or CDMA method to mitigate the near-far problem. In some DSSS implementations, the length of the DSSS spreading code may be varied according to the relative audibility of devices in a room. As described above with reference to FIG. 6, given the same amount of energy in the acoustic DSSS signal, if the spreading code increases the bandwidth of the acoustic DSSS signal, the acoustic DSSS signal will have a relatively lower maximum power and be relatively less audible. Alternatively or additionally, in some implementations, the calibration signals may be positioned orthogonal to one another. Some such implementations allow a system to simultaneously have DSSS signals with different spreading code lengths. Alternatively or additionally, in some implementations, the energy in each calibration signal may be modified to reduce the effects of the near-far problem (e.g., to boost the level of acoustic calibration signals generated by relatively quieter and / or more distant transmitting audio devices) and / or to obtain an optimal signal-to-noise ratio for a given operational purpose.
[0192] Figure 12 is a graph showing another example of a governance method. The elements of Figure 12 are as follows: 1210, 1211, 1212: frequency bands that do not overlap with each other; 230Ai, Bi, and Ci: multiple acoustic calibration signals time-domain multiplexed within frequency band 1210. While audio devices 1, 2, and 3 may appear to be using different portions of frequency band 1210, in this example, acoustic calibration signals 230Ai, Bi, and Ci span most or all of frequency band 1210; 230D and E: Multiple acoustic calibration signals code-domain multiplexed within frequency band 1211. While audio devices 4 and 5 may appear to use different portions of frequency band 1211, in this example acoustic calibration signals 230D and 230E span most or all of frequency band 1211; 230Aii, Bii, and Cii: Multiple acoustic calibration signals code-domain multiplexed within frequency band 1212. While audio devices 1, 2, and 3 may appear to use different portions of frequency band 1210, in this example, acoustic calibration signals 230Aii, Bii, and Cii span most or all of frequency band 1212.
[0193] 12 shows an example of how TDMA, FDMA, and CDMA may be used together in certain implementations of the present invention. In frequency band 1 (1210), TDMA is used to govern acoustic calibration signals 230Ai, Bi, and Ci transmitted by audio devices 1 through 3, respectively. Frequency band 1210 is a single frequency band, and acoustic calibration signals 230Ai, Bi, and Ci cannot fit within it simultaneously without overlapping.
[0194] In frequency band 2 (1211), CDMA is used to direct acoustic calibration signals 230D and 230E from audio devices 4 and 5, respectively. In this particular example, acoustic calibration signal 230D is longer in time than acoustic calibration signal 230E. A shorter calibration signal duration for audio device 5 may be useful if audio device 5 is louder than audio device 4, as, from the perspective of the receiving audio device, a shorter calibration signal duration corresponds to an increased bandwidth and lower peak frequency of the calibration signal. The signal-to-noise ratio (SNR) may also be improved with the relatively longer duration of acoustic calibration signal 230D.
[0195] In frequency band 3 (1212), CDMA is used to coordinate acoustic calibration signals 230Aii, Bii, and Cii transmitted by audio devices 1-3, respectively. These acoustic calibration signals correspond to alternative calibration signals used by audio devices 1-3, which are simultaneously transmitting TDMA-coordinated acoustic calibration signals for the same audio device within frequency band 1210. This is a form of FDMA in which longer calibration signals are placed in one frequency band (1212) and transmitted simultaneously (no TDMA), while shorter calibration signals are placed in another frequency band (1210) where TDMA is used.
[0196] FIG. 13 is a graph illustrating another example of a governance method. In this implementation, audio device 4 transmits acoustic calibration signals 230Di and 230Dii that are orthogonal to each other, and audio device 5 transmits acoustic calibration signals 230Ei and 230Eii that are also orthogonal to each other. In this example, all acoustic calibration signals are transmitted simultaneously within a single frequency band 1310. In this example, quadrature acoustic calibration signals 230Di and 230Ei are longer than in-phase calibration signals 230Dii and 230Eii transmitted by the two audio devices. As a result, each audio device has a faster, noisier set of observations derived from acoustic calibration signals 230Dii and 230Eii, albeit at a lower update rate, in addition to a higher SNR set of observations derived from acoustic calibration signals 230Di and 230Ei. This is an example of a CDMA-based governance method in which two audio devices are transmitting acoustic calibration signals designed for the acoustic space they share. In some cases, the governance method may also be based at least in part on the current listening objectives.
[0197] FIG. 14 illustrates elements of another example audio environment. In this example, audio environment 1401 is a multi-room residence including acoustic spaces 130A, 130B, and 130C. According to this example, doors 1400A and 1400B can alter the coupling of each acoustic space. For example, when door 1400A is open, acoustic spaces 130A and 130C are acoustically coupled to at least some extent, while when door 1400A is closed, acoustic spaces 130A and 130C are not acoustically coupled to any significant degree. In some implementations, a leader device may be configured to detect that a door has been opened (or another acoustic obstruction has been moved) according to the detection, or lack thereof, of audio device playback in an adjacent acoustic space.
[0198] In some examples, the coordinating device may coordinate all of the audio devices 100A-100E in all of the acoustic spaces 130A, 130B, and 130C. However, due to the significant level of acoustic isolation between the acoustic spaces 130A, 130B, and 130C when the doors 1400A and 1400B are closed, the coordinating device may, in some examples, treat the acoustic spaces 130A, 130B, and 130C as independent when the doors 1400A and 1400B are closed. In some examples, the coordinating device may treat the acoustic spaces 130A, 130B, and 130C as independent even when the doors 1400A and 1400B are open. However, in some instances, the coordinating device may manage audio devices located near the doors 1400A and / or 1400B, such that when the acoustic spaces are combined due to door opening, the audio devices near the open doors are treated as audio devices corresponding to the rooms on either side of the doors. For example, if the command device determines that door 1400A is open, the command device may be configured to consider audio device 100C to be an audio device for acoustic space 130A and also an audio device for acoustic space 130C.
[0199] FIG. 15 is a flow diagram outlining another example of the disclosed audio device governance method. The blocks of method 1500, as with other methods described herein, are not necessarily performed in the order shown. Furthermore, such methods may include more or fewer blocks than those shown and / or described. Method 1500 may be performed by a system including a governing device and governed audio devices. The system may include instances of apparatus 150 shown in FIG. 1B and described above, one of which is configured as the governing device. The governing device may, in some examples, include an instance of the governing module 213 disclosed herein.
[0200] According to this example, block 1505 involves steady-state operation of all participating audio devices. In this context, "steady-state" operation means operation according to the set of calibration signal parameters most recently received from the leading device. According to some implementations, the set of parameters may include one or more DSSS spreading code parameters and one or more DSSS carrier parameters.
[0201] In this example, block 1505 also involves one or more devices waiting for a trigger condition. The trigger condition can be, for example, an acoustic change in the audio environment in which the coordinated audio device is located. The acoustic change can be or include noise from a noise source, a change corresponding to a door or window being opened or closed (e.g., increased or decreased audibility of sound being played from one or more loudspeakers in an adjacent room), detected movement of an audio device in the audio environment, detected movement of a person in the audio environment, detected utterance of a person in the audio environment (e.g., utterance of a wake word), start of audio content playback (e.g., start of a movie, television program, music content, etc.), change in audio content playback (e.g., a volume change equal to or greater than a threshold change in decibels), etc. In some cases, the acoustic change is detected via an acoustic calibration signal (e.g., one or more acoustic scene metrics 225A estimated by baseband processor 218 of an audio device in the audio environment), for example, as disclosed herein.
[0202] In some instances, the trigger condition may be an indication that a new audio device has been powered on in the audio environment. In some such examples, the new audio device may be configured to generate one or more characteristic sounds that may or may not be audible to humans. According to some examples, the new audio device may be configured to play an acoustic calibration signal reserved for the new device.
[0203] In this example, in block 1510, it is determined whether a trigger condition is detected. If so, the process proceeds to block 1515. If not, the process returns to block 1505. In some implementations, block 1505 may include block 1510.
[0204] According to this example, block 1515 involves determining, by the governing device, one or more updated acoustic calibration signal parameters for one or more (in some cases all) of the governed audio devices and providing the updated acoustic calibration signal parameter(s) to the governed audio device(s). In some examples, block 1515 may involve providing, by the governing device, calibration signal information 205, described elsewhere herein. Determining the updated acoustic calibration signal parameters may involve using existing knowledge and estimates of the acoustic space, such as: · Device location; · Device range; Device orientation and relative angle of incidence; Relative clock bias and skew between devices; · Relative audibility of the device; · Indoor noise estimate; · Number of microphones and speakers on each device; -Directivity of each device's speakers; - Microphone directionality for each device; · The type of content being rendered in the acoustic space; the position of one or more listeners within the acoustic space; and / or ·Knowledge of acoustic space, including specular reflection and occlusion.
[0205] Such factors may, in some examples, be combined with operational objectives to determine a new operating point. Note that many of the parameters used as prior knowledge in determining updated calibration signal parameters may be derived from the acoustic calibration signal. Thus, it can be readily seen that a governed system, in some examples, can iteratively improve its performance as the system acquires more information, more accurate information, etc.
[0206] In this example, block 1520 includes reconfiguring, by one or more governed audio devices, one or more parameters used to generate the acoustic calibration signal according to the updated acoustic calibration signal parameter(s) received from the governing device. According to this implementation, after block 1520 is completed, the process returns to block 1505. Although an end is not shown in the flow diagram of FIG. 15, method 1500 can end in various ways, such as when an audio device is powered off.
[0207] FIG. 16 shows another example of an audio environment. The audio environment 130 shown in FIG. 16 is the same as that shown in FIG. 8, but also shows the angular separation of audio device 100B from the angular separation of audio device 100C from the perspective of audio device 100A (relative to audio device 100A). In FIG. 16, audio devices 100B and 100C are separated from device 100A by distances 810 and 811, respectively. In this particular situation, distance 811 is greater than distance 810. Assuming that audio devices 100B and 100C are producing audio device-playback sounds at approximately the same level, this means that audio device 100A will receive the acoustic calibration signal from audio device 100C at a lower level than the acoustic calibration signal from audio device 100B due to the additional acoustic loss caused by the longer distance 811.
[0208] This example focuses on the coordination of devices 100B and 100C to optimize device 100A's ability to hear both devices 100B and 100C. While there are other factors to consider, as outlined above, this example focuses on angle-of-arrival diversity caused by the angular separation of audio device 100B from audio device 100C relative to audio device 100A. Due to the difference in distances 810 and 811, the coordination may lead to longer code lengths for audio devices 100B and 100C to mitigate the near-far problem by reducing cross-channel correlation. However, when the receive beamformer (215) is implemented by audio device 100A, the angular separation between audio devices 100B and 100C places microphone signals corresponding to sounds from audio devices 100B and 100C in different lobes, providing further separation of the two received signals, thereby mitigating the near-far problem somewhat. This additional isolation may therefore allow the governing device to reduce the acoustic calibration signal length and acquire observations at a faster rate.
[0209] This not only applies to, for example, acoustic DSSS spreading code lengths: when spatial microphone feeds are used by audio device 100A (and / or audio devices 100B and 100C) instead of omnidirectional microphone feeds, any acoustic calibration parameters that can be changed to mitigate the near-far problem (e.g., even when using FDMA or TDMA) may no longer be necessary.
[0210] Control over spatial measures (in this case, angular diversity) depends on estimates of these characteristics already being available. In one example, calibration parameters may be optimized for the omnidirectional microphone feed (206), and then, after DoA estimates are available, acoustic calibration parameters may be optimized for the spatial microphone feed. This is one implementation of the trigger condition described above with reference to FIG. 15.
[0211] FIG. 17 is a block diagram illustrating example calibration signal demodulator elements, baseband processor elements, and calibration signal generator elements according to some disclosed implementations. As with other figures provided herein, the types and number of elements shown in FIG. 17 are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements. Other examples may implement other methods, such as frequency-domain correlation. In this example, calibration signal demodulator 214, baseband processor 218, and calibration signal generator 212 are implemented by instances of control system 160 described above with reference to FIG. 1B.
[0212] According to some implementations, there is one instance of calibration signal demodulator 214, baseband processor 218, and calibration signal generator 212 for each transmitted (played) acoustic calibration signal from each audio device for which the acoustic calibration signal is received. In other words, for the implementation shown in Figure 16, audio device 100A implements one instance of calibration signal demodulator 214, baseband processor 218, and calibration signal generator 212 corresponding to the acoustic calibration signal received from audio device 100B, and one instance of calibration signal demodulator 214, baseband processor 218, and calibration signal generator 212 corresponding to the acoustic calibration signal received from audio device 100C.
[0213] For purposes of illustration, the following description of Figure 17 will continue to use this example of audio device 100A of Figure 16 as the local device, i.e., as implementing instances of calibration signal demodulator 214, baseband processor 218, and calibration signal generator 212 in this example. More specifically, the following description of Figure 17 will assume that microphone signal 206 received by calibration signal demodulator 214 includes a reproduced sound produced by a loudspeaker of audio device 100B that includes an acoustic calibration signal generated by audio device 100B, and that the instances of calibration signal demodulator 214, baseband processor 218, and calibration signal generator 212 shown in Figure 17 correspond to the acoustic calibration signal reproduced by the loudspeaker of audio device 100B.
[0214] In this particular implementation, the calibration signal is a DSSS signal. Thus, according to this implementation, calibration signal generator 212 includes acoustic DSSS carrier module 1715 configured to provide calibration signal demodulator 214 with DSSS carrier replica 1705 of the DSSS carrier used by audio device 100B to generate the acoustic DSSS signal. In some alternative implementations, acoustic DSSS carrier module 1715 may be configured to provide calibration signal demodulator 214 with one or more DSSS carrier parameters used by audio device 100B to generate the acoustic DSSS signal. In some alternatives, the calibration signal is another type of calibration signal generated by modulating a carrier, such as a maximal length sequence or other type of pseudo-random binary sequence.
[0215] In this implementation, calibration signal generator 212 also includes an acoustic DSSS spreading code module 1720 configured to provide calibration signal demodulator 214 with a DSSS spreading code 1706 used by audio device 100B to generate the acoustic DSSS signal. DSSS spreading code 1706 corresponds to spreading code C(t) in the equations disclosed herein. DSSS spreading code 1706 may be, for example, a pseudo-random number (PRN) sequence.
[0216] According to this implementation, calibration signal demodulator 214 includes a bandpass filter 1703 configured to generate a bandpass filtered microphone signal 1704 from received microphone signal 206. In some instances, the passband of bandpass filter 1703 may be centered on the center frequency of the acoustic DSSS signal from audio device 100B being processed by calibration signal demodulator 214. Passband filter 1703 may, for example, pass a main lobe of the acoustic DSSS signal. In some examples, the passband of passband filter 1703 may be equal to the frequency band for transmission of the acoustic DSSS signal from audio device 100B.
[0217] In this example, calibration signal demodulator 214 includes multiplication block 1711A configured to convolve bandpass filtered microphone signal 1704 with DSSS carrier replica 1705 to generate baseband signal 1700. According to this implementation, calibration signal demodulator 214 also includes multiplication block 1711B configured to apply DSSS spreading code 1706 to baseband signal 1700 to generate de-spread baseband signal 1701.
[0218] According to this example, calibration signal demodulator 214 includes accumulator 1710A, and baseband processor 218 includes accumulator 1710B. Accumulators 1710A and 1710B are sometimes referred to herein as summing elements. Accumulator 1710A operates for a time period sometimes referred to herein as a “coherent time” that corresponds to the code length for each acoustic calibration signal (in this example, the code length for the acoustic DSSS signal currently being played by audio device 100B). In this example, accumulator 1710A implements an “integrate and dump” process; in other words, after summing despread baseband signal 1701 over the coherent time, accumulator 1710A outputs (“dumps”) demodulated coherent baseband signal 208 to baseband processor 218. In some implementations, demodulated coherent baseband signal 208 may be a single number.
[0219] In this example, the baseband processor 218 includes a square law module 1712, which in this example is configured to square the modulus of the demodulated coherent baseband signal 208 and output a power signal 1722 to an accumulator 1710B. After the modulus and squaring process, the power signal may be considered an incoherent signal. In this example, the accumulator 1710B operates over an "incoherent time." The incoherent time may, in some examples, be based on an input from a governing device. The incoherent time may, in some examples, be based on a desired SNR. According to this example, the accumulator 1710B outputs a delayed waveform 400 at multiple delays (also referred to herein as "tau" or instances of tau (τ)).
[0220] Steps 1704 to 208 in FIG. 17 can be expressed as follows:
number
[0221] In the above equation, Y(tau) represents the coherent demodulator output (208), d[n] represents the bandpass filtered signal (1704 or A in FIG. 17), CA represents the local copy of the spreading code used to modulate the calibration signal (in this example, a DSSS signal) by the far-field device in the room (in this example, audio device 100B), and the last term is the carrier signal. In some examples, all of these signal parameters may be coordinated between audio devices in the audio environment (e.g., determined and provided by a coordinated device).
[0222] From Y(tau)(208)<Y(tau)> The signal chain in FIG. 17 to (400) is an incoherent integration in which the coherent demodulator output is squared and averaged. The number of averages (the number of times the incoherent accumulator 1710B runs) is a parameter that may be determined and provided by a governing device, in some instances, based on, for example, a determination that a sufficient SNR has been achieved. In some instances, an audio device implementing the baseband processor 218 may determine the number of averages, for example, based on a determination that a sufficient SNR has been achieved.
[0223] Incoherent integration can be expressed mathematically as follows:
number
[0224] The above equation involves simply averaging the squared coherent delayed waveform over a time period defined by N, where N represents the number of blocks used in the incoherent integration.
[0225] 18 illustrates elements of a calibration signal demodulator according to another example. According to this example, calibration signal demodulator 214 is configured to generate a delay estimate, a DoA estimate, and an audibility estimate. In this example, calibration signal demodulator 214 is configured to perform coherent demodulation, and then incoherent integration is performed on the fully delayed waveform. As with the example described above with reference to FIG. 17, this example assumes that calibration signal demodulator 214 is implemented by audio device 100A and configured to demodulate an acoustic DSSS signal played by audio device 100B.
[0226] In this example, calibration signal demodulator 214 includes bandpass filter 1703 configured to remove unwanted energy from other audio signals, such as the portion of the audio content being rendered for the listener's experience and the acoustic DSSS signal placed in other frequency bands to avoid near-far problems. For example, bandpass filter 1703 may be configured to pass energy from one of the frequency bands shown in Figures 12 and 13.
[0227] Matched filter 1811 is configured to calculate delayed waveform 1802 by correlating bandpass filtered signal 1704 with a local replica of the acoustic calibration signal of interest. In this example, the local replica is an instance of DSSS signal replica 204 corresponding to the DSSS signal generated by audio device 100B. Matched filter output 1802 is then lowpass filtered by lowpass filter 712 to generate coherently demodulated complex delayed waveform 208. In some alternative implementations, lowpass filter 712 may be located after the squaring operation in baseband processor 218 to generate the incoherently averaged delayed waveform, such as the example described above with reference to FIG. 17.
[0228] In this example, the channel selector 1813 is configured to control the bandpass filter 1703 (e.g., the passband of the bandpass filter 1703) and the matched filter 1811 according to the calibration signal information 205. As described above, the calibration signal information 205 may include parameters used by the control system 160 to demodulate the calibration signal, etc. The calibration signal information 205 may, in some examples, indicate which audio device is generating the acoustic calibration signal. In some examples, the calibration signal information 205 may be received (e.g., via wireless communication) from an external source, such as a command device.
[0229] FIG. 19 is a block diagram illustrating an example of a baseband processor element according to some disclosed implementations. As with other figures provided herein, the types and number of elements shown in FIG. 19 are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements. In this example, baseband processor 218 is implemented by an instance of control system 160 described above with reference to FIG. 1B.
[0230] In this particular implementation, no coherent techniques are applied. Thus, the first operation performed is to take the power of the complex delayed waveform 208 via the square law module 1712 to generate the incoherent delayed waveform 1922. The incoherent delayed waveform 1922 is integrated by the accumulator 1710B over a period of time (which in this example is specified in the calibration signal information 205 received from the governing device, but may be determined locally in some examples) to generate the incoherently averaged delayed waveform 400. According to this example, the delayed waveform 400 is then processed in several ways, as follows: 1. The leading edge estimator 1912 is configured to obtain a delay estimate 1902. The delay estimate 1902 is an estimated time delay of the received signal. In some examples, the delay estimate 1902 may be based at least in part on an estimate of the position of the leading edge of the delayed waveform 400. According to some such examples, the delay estimate 1902 may be determined according to the number of time samples in the signal portion (e.g., positive portion) of the delayed waveform up to and including the time sample corresponding to the position of the leading edge of the delayed waveform 400 or a time sample that is less than one chip period (inversely proportional to the signal bandwidth) after the position of the leading edge of the delayed waveform 400. In the latter case, according to some examples, this delay may be used to compensate for the width of the autocorrelation of the DSSS code. As the chipping rate increases, the width of the autocorrelation peak narrows and is minimized when the chipping rate is equal to the sampling rate. This condition (chipping rate equal to sampling rate) results in a delayed waveform 400 that is the closest approximation to the true impulse response for the audio environment for a given DSSS code. As the chipping rate increases, spectral overlap (aliasing) may occur following calibration signal modulator 220A. In some examples, calibration signal modulator 220A may be bypassed or omitted when the chipping rate is equal to the sampling rate. A chipping rate that approaches that of the sampling rate (e.g., a chipping rate that is 80% of the sampling rate, 90% of the sampling rate, etc.) may provide a delayed waveform 400 that is a satisfactory approximation of the actual impulse response for some purposes. In some such examples, delay estimate 1902 may be based in part on information about calibration signal characteristics (e.g., DSSS signal characteristics). In some examples, leading edge estimator 1912 may be configured to estimate the position of the leading edge of delayed waveform 400 according to the first instance of a value greater than a threshold during a time window. Some examples are described below with reference to FIG. 20.In another example, the leading edge estimator 1912 may be configured to estimate the position of the leading edge of the delayed waveform 400 according to the position of a maximum value (e.g., a maximum value within a time window), which is an example of “peak-picking.” Note that many other techniques can be used to estimate the delay (e.g., peak-picking). 2. In this example, the baseband processor 218 is configured to perform DoA estimation 1903 by windowing the delayed waveform 400 (using windowing block 1913) before using the delay-and-sum DoA estimator 1914. The delay-and-sum DoA estimator 1914 may perform DoA estimation based at least in part on a determination of the steered response power (SRP) of the delayed waveform 400. Thus, the delay-and-sum DoA estimator 1914 may also be referred to herein as an SRP module or delay-and-sum beamformer. Windowing is useful for isolating a time interval around a leading edge, so that the resulting DoA estimate is based more on signal than noise. In some examples, the window size may be in the range of tens or hundreds of milliseconds, e.g., 10 to 200 milliseconds. In some cases, the window size may be selected based on knowledge of typical room decay times or based on knowledge of the decay times of the audio environment in question. In some cases, the window size may be adaptively updated over time. For example, some implementations may involve determining a window size that results in at least some portion of the window being occupied by a signal portion of the delayed waveform 400. Some such implementations may involve estimating noise power according to time samples occurring before the leading edge. Some such implementations may involve selecting a window size that results in at least a threshold percentage of the window being occupied by a portion of the delayed waveform that corresponds to at least a threshold signal level, e.g., at least 6 dB greater than the estimated noise power, at least 8 dB greater than the estimated noise power, at least 10 dB greater than the estimated noise power, etc. 3. According to this example, the baseband processor 218 is configured to perform the audibility estimation 1904 by estimating signal-to-noise power using the SNR estimation block 1915. In this example, the SNR estimation block 1915 is configured to extract the signal power estimate 402 and the noise power estimate 401 from the delayed waveform 400. According to some such examples, the SNR estimation block 1915 may be configured to determine the signal portion and the noise portion of the delayed waveform 400, as described below with reference to FIG. 20 . In some such examples, the SNR estimation block 1915 may be configured to determine the signal power estimate 402 and the noise power estimate 401 by averaging the signal portion and the noise portion over a selected time window. In some such examples, the SNR estimation block 1915 may be configured to perform the SNR estimation according to the ratio of the signal power estimate 402 to the noise power estimate 401. In some instances, the baseband processor 218 may be configured to perform the audibility estimation 1904 according to the SNR estimate. For a given amount of noise power, the SNR is proportional to the audibility of an audio device. Thus, in some implementations, the SNR may be used directly as a proxy for (e.g., a value proportional to) an estimate of the actual audio device audibility. Some implementations involving calibrated microphone feeds may involve measuring absolute audibility (e.g., in dBSPL) and converting the SNR to an absolute audibility estimate. In some such implementations, the method for determining the absolute audibility estimate takes into account acoustic losses due to distance between audio devices and the variability of noise within the room. In other implementations, there are other techniques for estimating signal power, noise power, and / or relative audibility from delayed waveforms.
[0231] 20 shows an example of a delay waveform. In this example, delay waveform 400 is output by an instance of baseband processor 218. According to this example, the vertical axis represents power and the horizontal axis represents pseudorange in meters. As mentioned above, baseband processor 218 is configured to extract delay information, sometimes referred to herein as τ, from the demodulated acoustic calibration signal. The value of τ can be converted to a pseudorange measurement, sometimes referred to herein as ρ, as follows: ρ=τc
[0232] In the above equation, c is the speed of sound. In Figure 20, delayed waveform 400 includes a noise portion 2001 (sometimes called the noise floor) and a signal portion 2002. Negative values in the pseudorange measurements (and corresponding delayed waveforms) can be identified as noise: since negative ranges (distances) have no physical meaning, the power corresponding to negative pseudoranges is assumed to be noise.
[0233] In this example, signal portion 2002 of waveform 400 includes a leading edge 2003 and a trailing edge. If the power of signal portion 2002 is relatively strong, leading edge 2003 is a prominent feature of delayed waveform 400. In some examples, leading edge estimator 1912 of FIG. 19 may be configured to estimate the position of leading edge 2003 according to the first instance of a power value greater than a threshold during a time window. In some examples, the time window may begin when τ (or ρ) is 0. In some cases, the window size may be in the range of tens or hundreds of milliseconds, e.g., in the range of 10 to 200 milliseconds. According to some implementations, the threshold may be a previously selected value, e.g., −5 dB, −4 dB, −3 dB, −2 dB, etc. In some alternative examples, the threshold may be based on the power in at least a portion of delayed waveform 400, e.g., the average power of a noise portion.
[0234] However, as described above, in other examples, the leading edge estimator 1912 may be configured to estimate the position of the leading edge 2003 according to the position of a maximum value (e.g., a maximum value within a time window). In some instances, the time window may be selected as described above.
[0235] 19 may, in some examples, be configured to determine an average noise value corresponding to at least a portion of noise portion 2001 and an average or peak signal value corresponding to at least a portion of signal portion 2002. SNR estimation block 1915 of FIG. 19 may, in some such examples, be configured to estimate the SNR by dividing the average signal value by the average noise value.
[0236] Noise compensation (e.g., automatic leveling of loudspeaker-played content) to compensate for environmental noise conditions is a well-known and desirable feature, but has not previously been optimally implemented. Using microphones to measure environmental noise conditions presents a major challenge for noise estimation (e.g., online noise estimation), which is required to also measure loudspeaker-played content and implement noise compensation.
[0237] Because people in an audio environment may generally be outside the critical acoustic distance of any given room, echoes introduced from other devices a similar distance away may still represent a significant echo effect. Even if sophisticated multi-channel echo cancellation were available and managed to achieve the required performance, the logistics of providing a remote echo reference to the canceller may have unacceptable bandwidth and complexity costs.
[0238] Some disclosed implementations provide a method for continuously calibrating a constellation of audio devices in an audio environment through persistent (e.g., continuous, or at least ongoing) characterization of the acoustic space, including people, devices, and audio conditions (such as noise and / or echo). In some disclosed examples, such a process continues even while media is being played through the audio devices in the audio environment.
[0239] As used herein, a "gap" in a playback signal refers to a time (or time interval) in the playback signal where playback content is missing (or has a level below a predetermined threshold). For example, a "gap" (also referred to herein as a "forced gap" or "parameterized forced gap") may be an attenuation of playback content in a certain frequency range during a certain time interval. In some disclosed implementations, gaps may be inserted in one or more frequency ranges of an audio playback signal of a content stream to generate a modified audio playback signal, which may be reproduced or "played back" in an audio environment. In some such implementations, N gaps may be inserted in N frequency ranges of the audio playback signal during N time intervals.
[0240] According to some such implementations, the M audio devices may coordinate their gaps in time and frequency, thereby allowing accurate detection of the far-field (for each device) at the gap frequency and time interval. These “orchestrated gaps” are an important aspect of this disclosure. In some examples, M may be a number corresponding to all audio devices in the audio environment. In some instances, M may be a number corresponding to all audio devices in the audio environment except for a target audio device, where the target audio device is an audio device whose played-back audio is sampled by one or more microphones of the M orchestrated audio devices in the audio environment (e.g., one or more microphones of the M orchestrated audio devices in the audio environment), e.g., to assess the relative audibility, location, nonlinearity, and / or other characteristics of the target audio device. In some examples, the target audio device may play an unmodified audio playback signal that does not include any inserted gaps in any frequency range. In other examples, M may be a number corresponding to a subset of the audio devices in the audio environment, e.g., multiple participating non-target audio devices.
[0241] It is desirable that the controlled gap should have a low perceptual impact (e.g., negligible perceptual impact) to a listener in the audio environment. Thus, in some examples, gap parameters may be selected to minimize the perceptual impact.
[0242] In some examples, while the modified audio playback signal is being played in the audio environment, the target device may be playing an unmodified audio playback signal that does not include gaps inserted in any frequency ranges. In such examples, the relative audibility and / or location of the target device may be estimated from the perspective of the M audio devices that are playing the modified audio playback signal.
[0243] Figure 21 shows another example of an audio environment. As with other figures provided herein, the types and number of elements shown in Figure 21 are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements.
[0244] According to this example, audio environment 2100 includes a primary living space 2101a and a room 2101b adjacent to primary living space 2101a, where a wall 2102 and a door 2111 separate primary living space 2101a from room 2101b. In this example, the amount of acoustic isolation between primary living space 2101a and room 2101b depends on whether door 2111 is open or closed, and, if open, how open it is.
[0245] At the time corresponding to Figure 21, a smart television (TV) 2103a is located within the audio environment 2100. According to this example, the smart TV 2103a includes a left speaker 2103b and a right speaker 2103c.
[0246] In this example, smart audio devices 2104, 2105, 2106, 2107, 2108, 2109, and 2113 are also located within audio environment 2100 at times corresponding to Figure 21. According to this example, smart audio devices 2104-2109 each include at least one microphone and at least one loudspeaker. However, in this example, smart audio devices 2104-2109 and 2113 include loudspeakers of different sizes and with different capabilities.
[0247] According to this example, at least one acoustic event occurs within the audio environment 2100. In this example, one acoustic event is caused by a speaker 2110 uttering a voice command 2112.
[0248] In this example, another acoustic event is caused, at least in part, by variable element 2115, which is a door in audio environment 2100. According to this example, when door 2115 is open, sounds from outside the environment may be perceived more clearly inside audio environment 2100. Furthermore, changing the angle of door 2115 changes some of the echo paths within audio environment 2100. According to this example, element 2114 represents a variable element in the impulse response of audio environment 2100 that is caused by changing the position of door 2115.
[0249] In some examples, a series of forced gaps are inserted into the playback signal, each in a different frequency band (or set of bands) of the playback signal, allowing a pervasive listener to monitor non-playback sound occurring “at” each forced gap. Here, occurring “at” a gap means occurring in the frequency band(s) in which the gap is inserted during the time interval in which the gap occurs. FIG. 22A is an example spectrogram of a modified audio playback signal. In this example, the modified audio playback signal was created by inserting gaps into an example audio playback signal. More specifically, to generate the spectrogram of FIG. 22A , the disclosed method was performed on the audio playback signal to introduce forced gaps into its frequency bands (e.g., gaps G1, G2, and G3 shown in FIG. 22A ), thereby generating the modified audio playback signal. In the spectrogram shown in FIG. 22A , position along the horizontal axis indicates time, and position along the vertical axis indicates the frequency of the content of the modified audio playback signal at a given moment. The density of dots in each small region (in this example, each such region is centered at a point having vertical and horizontal coordinates) indicates the energy of the content of the modified audio playback signal at the corresponding frequency and time: higher density regions indicate content with greater energy, and lower density regions indicate content with less energy. Thus, gap G1 occurs at an earlier time (i.e., time interval) than gaps G2 or G3 occur, and gap G1 is inserted in a higher frequency band than the frequency bands into which gaps G2 or G3 are inserted.
[0250] The introduction of forced gaps into a playback signal according to some disclosed methods differs from simplex device behavior, in which a device pauses a playback stream of content (e.g., to better hear the user and the user's environment). The introduction of forced gaps into a playback signal according to some disclosed methods may be optimized to significantly reduce (or eliminate) the perceptibility of artifacts resulting from the introduced gaps during playback, preferably such that the forced gaps have no or minimal perceptible impact on the user, but the output signal of a microphone in the playback environment indicates the forced gaps (e.g., so that the gaps can be utilized to implement pervasive listening methods). By using forced gaps introduced according to some disclosed methods, a pervasive listening system can monitor non-playback sounds (e.g., sounds indicative of background activity and / or noise in the playback environment) without the use of an acoustic echo canceller.
[0251] Referring to FIGS. 22B and 22C, we will now describe examples of parameterized forced gaps that may be inserted in frequency bands of an audio playback signal, and criteria for selecting the parameters of such forced gaps. FIG. 22B is a graph illustrating an example of a gap in the frequency domain. FIG. 22C is a graph illustrating an example of a gap in the time domain. In these examples, the parameterized forced gap is an attenuation of the playback content using a band attenuation G, whose profile over both time and frequency resembles the profiles shown in FIGS. 22B and 22C. Here, the gap is forced by applying an attenuation G to the playback signal over a range of frequencies (a "band") defined by a center frequency f0 (shown in FIG. 22B) and a bandwidth B (also shown in FIG. 22B). Here, the attenuation varies as a function of time at each frequency in the frequency band (e.g., at each frequency bin within the frequency band) with a profile similar to that shown in FIG. 22C. The maximum value of attenuation G (as a function of frequency across the band) can be controlled to increase from 0 dB (at the lowest frequency of the band) to a maximum attenuation (suppression depth) Z at a center frequency f0 (shown in FIG. 22B), and then decrease (with increasing frequency above the center frequency) to 0 dB (at the highest frequency of the band).
[0252] In this example, the graph in Figure 22B shows a profile of band attenuation G as a function of frequency (i.e., frequency bins) that is applied to frequency components of an audio signal to force gaps in the audio content of the signal within the band. The audio signal may be a playback signal (e.g., a channel of a multi-channel playback signal), and the audio content may be playback content.
[0253] According to this example, the graph in FIG. 22C shows the profile of band attenuation G as a function of time applied to the frequency component at center frequency f0 to force the gap shown in FIG. 22B into the audio content of the signal in the band. For each other frequency component in the band, the band gain as a function of time may have a profile similar to that shown in FIG. 22C, but the suppression depth Z in FIG. 22C may be replaced by an interpolated suppression depth kZ, where k is a factor ranging from 0 to 1 (as a function of frequency) in this example, such that kZ has the profile shown in FIG. 22B. In some examples, for each frequency component, the attenuation G may also be interpolated (as a function of frequency) from 0 dB to suppression depth kZ (e.g., at the center frequency, k = 1, as shown in FIG. 22C). This may be to reduce musical artifacts resulting from the introduction of gaps, for example. Three regions (time intervals) t1, t2, and t3 of this latter interpolation are shown in FIG. 22C.
[0254] Thus, when a gap forcing operation is performed for a particular frequency band (e.g., a band centered around center frequency f0 as shown in FIG. 22B), the attenuation G applied to each frequency component within the band (e.g., each bin within the band) in this example follows the trajectory shown in FIG. 22C. It starts at 0 dB, drops to a depth of -kZ dB at t1 seconds, remains there for t2 seconds, and finally rises back to 0 dB at t3 seconds. In some implementations, the total time t1 + t2 + t3 may be selected taking into account the time resolution of any frequency translation being used to analyze the microphone feed, as well as a reasonable duration that is not too intrusive to the user. Some examples of t1, t2, and t3 for single-device implementations are shown in Table 1 below.
[0255] Some disclosed methods cover the entire frequency spectrum of the audio playback signal, count bands (where B count is a number, for example, B countThis involves inserting forced gaps according to a predetermined fixed banding structure, including a band width (where σ is σ = 49). To force a gap in one of the bands, in such an example, a band attenuation is applied in the band. Specifically, for the jth band, an attenuation G j may be applied over the frequency range defined by that band.
[0256] Table 1 below shows exemplary values for the parameters t1, t2, t3, the depth Z for each band, and the number of bands B for a single device implementation. count Here is an example: [Table 1]
[0257] In determining the number of bands and the width of each band, there is a trade-off between the perceptual impact of the gaps and their usefulness: narrower bands with gaps are better in that they typically have less perceptual impact, whereas wider bands with gaps are better for implementing noise estimation (and other pervasive listening methods) in all frequency bands across the entire frequency spectrum (e.g., in response to changes in background noise or playback environment conditions) and reducing the time required ("convergence" time) to converge to a new noise estimate (or other value monitored by pervasive listening). If only a limited number of gaps can be forced at one time, sequentially forcing gaps in many small bands takes longer than sequentially forcing gaps in fewer, larger bands, leading to relatively longer convergence times. Larger bands (with gaps) provide more information about the background noise (or other value monitored by pervasive listening) at one time, but generally have a greater perceptual impact.
[0258] Our initial work introduced gaps in single-device situations where echo effects were primarily (or entirely) near-field. Near-field echoes are significantly affected by the direct path of audio from the speaker to the microphone. This characteristic applies to almost all compact duplex audio devices (such as smart audio devices), with the exception of devices with larger enclosures and significant acoustic decoupling. By introducing a short, perceptually masked gap in playback, as shown in Table 1, an audio device can learn about one edge of the acoustic space in which it is deployed through its own echo.
[0259] However, when other audio devices are also playing content in the same audio environment, the inventors have discovered that the usefulness of a single audio device's gaps is reduced due to far-field echo corruption. Far-field echo corruption often degrades the performance of local echo cancellation, significantly degrading overall system performance. Far-field echo corruption is difficult to remove for a variety of reasons. One reason is that acquiring a reference signal can require increased network bandwidth and added complexity due to additional delay estimation. Furthermore, as noise conditions increase and responses become longer (more reverberant and spread out in time), estimating the far-field impulse response becomes more difficult. In addition, far-field echo corruption is typically correlated with near-field echoes and other far-field echo sources, further complicating far-field impulse response estimation.
[0260] The inventors have discovered that when multiple audio devices in an audio environment coordinate their gaps in time and frequency, a clearer perception of the far field (for each audio device) can be obtained when the multiple audio devices play modified audio playback signals. The inventors have also discovered that when multiple audio devices play modified audio playback signals and a target audio device plays an unmodified audio playback signal, the relative audibility and location of the target device can be estimated from the perspective of each of the multiple audio devices, even while media content is being played.
[0261] Furthermore, perhaps counterintuitively, the inventors have discovered that breaking guidelines previously used for single-device implementations (e.g., leaving the gap open for a longer period of time than shown in Table 1) leads to an implementation suitable for multiple devices making cooperative measurements via coordinated gaps.
[0262] For example, in some coordinated gap implementations, t2 may be longer than shown in Table 1 to accommodate varying acoustic path lengths (acoustic delays) between multiple distributed devices in an audio environment, which may be on the order of meters (as opposed to a fixed microphone-speaker acoustic path length in a single device, which may be at most tens of centimeters). In some examples, the default t2 value may be 25 milliseconds larger than the 80 millisecond value shown in Table 1 to accommodate separations of up to 8 meters between coordinated audio devices. In some coordinated gap implementations, the default t2 value may be longer than the 80 millisecond value shown in Table 1 for another reason: in a coordinated gap implementation, a longer t2 is preferred to accommodate timing misalignment of coordinated audio devices to ensure that a sufficient amount of time has passed during which all coordinated audio devices have reached their Z attenuation values. In some examples, an additional 5 milliseconds may be added to the default value of t2 to accommodate timing misalignment. Thus, in some governed gap implementations, the default value for t2 may be 110 ms, with a minimum value of 70 ms and a maximum value of 150 ms.
[0263] In some coordinated gap implementations, t1 and / or t3 may also differ from the values shown in Table 1. In some examples, t1 and / or t3 may be adjusted as a result of a listener's inability to perceive the different times at which devices enter or exit their decay periods due to timing issues and physical distance discrepancies. Due, at least in part, to spatial masking (resulting from multiple devices playing audio from different locations), a listener's ability to perceive the different times at which coordinated audio devices enter or exit their decay periods tends to be lower than in a single-device scenario. Therefore, in some coordinated gap implementations, the minimum values of t1 and t3 may be reduced and the maximum values of t1 and t3 may be increased compared to the single-device example shown in Table 1. According to some such examples, the minimum values of t1 and t3 may be reduced to 2, 3, or 4 milliseconds, and the maximum values of t1 and t3 may be increased to 20, 25, or 30 milliseconds.
[0264] Examples of measurements using the governed gap FIG. 22D shows an example of a modified audio playback signal including coordinated gaps for multiple audio devices in an audio environment. In this implementation, multiple smart devices in an audio environment coordinate gaps to estimate their relative audibility with one another. In this example, one measurement session corresponding to one gap is conducted during a time interval, and the measurement session includes only devices in the primary living space 2101a of FIG. 21. According to this example, previous audibility data indicates that smart audio device 2109 located in room 2101b has already been classified as barely audible to other audio devices and is located in a separate zone.
[0265] In the example shown in FIG. 22D, the governed gap is k where k represents the center frequency of the frequency band being measured. The elements shown in Figure 22D are: Graph 2203 shows the G in dB for smart audio device 2113 in Figure 21. k is the plot of; Graph 2204 shows the Gamma in dB for smart audio device 2113 of Figure 21. k is the plot of; Graph 2205 shows the G in dB for the smart audio device 2113 in Figure 21. k is the plot of; Graph 2206 shows the G in dB for smart audio device 2113 in Figure 21. k is the plot of; Graph 2207 shows the G in dB for the smart audio device 2113 in Figure 21. k is the plot of; Graph 2208 shows the Gamma in dB for smart audio device 2113 of Figure 21. k is the plot of; Graph 2209 shows the G in dB for the smart audio device 2113 in Figure 21. k This is the plot.
[0266] As used herein, the term "session" (also referred to herein as "measurement session") refers to a period of time during which frequency range measurements are performed. During a measurement session, a set of frequencies with associated bandwidths, as well as a set of participating audio devices, may be specified.
[0267] One audio device may optionally be designated as the "target" audio device for a measurement session. When the target audio device participates in the measurement session, according to some examples, the target audio device may be permitted to ignore forced gaps and play the unmodified audio playback signal during the measurement session. According to some such examples, other participating audio devices may hear the target device playback, including the target device playback within the frequency range being measured.
[0268] As used herein, the term "audibility" refers to the degree to which a device can hear the speaker output of another device. Some examples of audibility are provided below.
[0269] According to the example shown in FIG. 22D , at time t1, the leader device initiates a measurement session with the target audio device, smart audio device 2113, and selects one or more bin center frequencies to be measured, including frequency k. In some examples, the leader device may be a smart audio device acting as a leader. In other examples, the leader device may be another leader device, such as a smart home hub. This measurement session runs from time t1 to time t2. The other participating smart audio devices, smart audio devices 2104-2108, apply gaps in their outputs and play modified audio playback signals, while smart audio device 2113 plays the unmodified audio playback signal.
[0270] The subset of smart audio devices in audio environment 2100 (smart audio devices 2104-2108) that are playing the modified audio playback signal including the coordinated gaps is one example of what may be referred to as M audio devices. According to this example, smart audio device 2109 also plays the unmodified audio playback signal. Therefore, smart audio device 2109 is not one of the M audio devices. However, because smart audio device 2109 is inaudible to other smart audio devices in the audio environment, smart audio device 2109 is not the target audio device in this example, despite the fact that both smart audio device 2109 and the target audio device (smart audio device 2113 in this example) play the unmodified audio playback signal.
[0271] It is desirable that the controlled gap should have a low perceptual impact (e.g., negligible perceptual impact) to a listener in the audio environment during a measurement session. Thus, in some examples, gap parameters may be selected to minimize the perceptual impact. Some examples are described below with reference to Figures 22B-22E.
[0272] During this time (measurement session from time t1 to time t2), smart audio devices 2104-2108 receive reference audio bins from the target audio device (smart audio device 2113) for time-frequency data for this measurement session. In this example, the reference audio bins correspond to the playback signal that smart audio device 2113 uses as a local reference for echo cancellation. Smart audio device 2113 has access to these reference audio bins for the purposes of audibility measurements as well as echo cancellation.
[0273] According to this example, at time t2, the first measurement session ends and the governing device begins a new measurement session, this time selecting one or more bin center frequencies that do not include frequency k. In the example shown in FIG. 22D, no gaps are applied for frequency k during the period t2 to t3, and thus the graph shows unity gain for all devices. In some such examples, the governing device may insert a series of gaps into each of multiple frequency ranges for a sequence of measurement sessions for bin center frequencies that do not include frequency k. For example, the governing device may insert second through Nth gaps into second through Nth frequency ranges of the audio playback signal during second through Nth time intervals for second through Nth subsequent measurement sessions while the smart audio device 2113 remains the target audio device.
[0274] In some such examples, the leader device may then select another target audio device, for example, smart audio device 2104. The leader device may instruct smart audio device 2113 to be one of the M smart audio devices playing the modified audio playback signal with the coordinated gap. The leader device may instruct the new target audio device to play the unmodified audio playback signal. According to some such examples, after the leader device has conducted N measurement sessions for the new target audio device, the leader device may select another target audio device. In some such examples, the leader device may continue to conduct measurement sessions until a measurement session has been performed for each participating audio device in the audio environment.
[0275] In the example shown in FIG. 22D , a different type of measurement session occurs between times t3 and t4. According to this example, at time t3, in response to user input (e.g., a voice command to a smart audio device acting as the leader device), the leader device initiates a new session to fully calibrate the loudspeaker setup of the audio environment 2100. In general, users may be relatively more tolerant of a leader gap, which has a relatively higher perceptual impact, during a “setup” or “recalibration” measurement session, such as that performed between times t3 and t4. Thus, in this example, a large, contiguous set of frequencies, including k, is selected for measurement. According to this example, smart audio device 2106 is selected as the first target audio device during this measurement session. Thus, during the first phase of the measurement session, from times t3 to t4, all smart audio devices except smart audio device 2106 apply the gap.
[0276] Gap Bandwidth Figure 23A is a graph showing an example of the filter responses used to create the gap and to measure the frequency domain of the microphone signal used during the measurement session. According to this example, the elements of Figure 23A are as follows: Element 2301 represents the magnitude response of the filter used to generate the gap in the output signal; Element 2302 represents the magnitude response of the filter used to measure the frequency region corresponding to the gap caused by element 2301; Elements 2303 and 2304 represent the −3 dB points of 2301 at frequencies f1 and f2; Elements 2305 and 2306 represent the -3 dB points of 2302 at frequencies f3 and f4.
[0277] The bandwidth (BW_gap) of the gap response 2301 may be found by taking the difference between the -3 dB points 2303 and 2304: BW_gap=f2-f1, and BW_measure (bandwidth of the measured response 2302)=f4-f3.
[0278] According to one example, the quality of the measurement can be expressed as:
number
[0279] Because the bandwidth of the measurement response is typically fixed, the quality of the measurement can be adjusted by increasing the bandwidth of the gap filter response (e.g., widening the bandwidth). However, the bandwidth of the introduced gap is proportional to its perceptibility. Therefore, the bandwidth of the gap filter response should generally be determined taking into account both the quality of the measurement and the perceptibility of the gap. Some example quality values are shown in Table 2. [Table 2]
[0280] Although Table 2 shows "minimum" and "maximum" values, these values are for the purposes of this example only. Other implementations may involve quality values lower than 1.5 and / or higher than 3.
[0281] Gap Allocation Strategy A gap may be defined by: · Fundamental division of the frequency spectrum using center frequencies and measurement bandwidths; Aggregation of these minimum measurement bandwidths in a structure called banding; · duration, attenuation depth, and inclusion of one or more consecutive frequencies that fit into said agreed division of said frequency spectrum; Other time behaviors, such as ramping the attenuation depth at the beginning and end of the gap.
[0282] According to some implementations, the gaps may be selected according to a strategy that aims to measure and observe as much of the audible spectrum as possible in the shortest possible time while satisfying applicable perceptibility constraints.
[0283] Figures 23B, 23C, 23D, and 23E are graphs illustrating example gap allocation strategies. In these examples, time is represented by distance along the horizontal axis, and frequency is represented by distance along the vertical axis. These graphs provide examples to illustrate the patterns produced by various gap allocation strategies and how long it takes to measure a complete audio spectrum. In these examples, each coordinated gap measurement session is 10 seconds in length. As with other disclosed implementations, these graphs are provided solely as examples. Other implementations may include more, fewer, and / or different types, numbers, and / or sequences of elements. For example, in other implementations, each coordinated gap measurement session may be longer or shorter than 10 seconds. In these examples, the unshaded regions 2310 (sometimes referred to herein as "tiles") of the time / frequency space depicted in Figures 23B-23E represent gaps in the indicated time-frequency period (10 seconds). The medium-shaded regions 2315 represent frequency tiles that have been measured at least once. The lightly shaded area 2320 has not yet been measured.
[0284] Assuming the task at hand requires participating audio devices to insert coordinated gaps to "listen through to the room" (e.g., to assess noise, echo, etc. in the audio environment), the measurement session completion time is as shown in Figures 23B-23E. If the task requires each audio device to be targeted in turn and heard by the other audio devices, the time must be multiplied by the number of audio devices participating in the process. For example, if each audio device is targeted in turn, the 3 minutes and 20 seconds (3m20s) shown as the measurement session completion time in Figure 23B means that a system of seven audio devices will be fully mapped after 7*3m20s=23m20s. When cycling through frequencies / bands and enforcing multiple gaps at once, in these examples the gaps are spaced as far apart in frequency as possible for efficiency in covering the spectrum.
[0285] Figures 23B and 23C are graphs illustrating examples of coordinated gap sequences according to a gap allocation strategy. In these examples, the gap allocation strategy involves gapping N entire frequency bands at a time (each frequency band includes at least one frequency bin, and in most cases multiple frequency bins) during each successive measurement session. In Figure 23B, N=1, and in Figure 23C, N=3, meaning that the example in Figure 23C involves inserting three gaps during the same time interval. In these examples, the banding structure used is a 20-band Mel-spaced arrangement. According to some such examples, the sequence may resume after all 20 frequency bands have been measured. While 3 m20 s is a reasonable time to achieve a complete measurement, the gaps punched in the critical audio region of 300 Hz to 8 kHz are very wide, resulting in a significant amount of time being spent measuring outside of this region. Due to the relatively wide gaps in the 300 Hz to 8 kHz frequency range, this particular strategy is highly perceptible to the user.
[0286] Figures 23D and 23E are graphs showing examples of sequences of enforced gaps according to another gap allocation strategy. In these examples, the gap allocation strategy involves modifying the banding structure shown in Figures 23B and 23C to map to an "optimized" frequency range of approximately 300 Hz to 8 kHz. While the sequence ends slightly earlier because the 20th band is ignored, the overall allocation strategy is otherwise unchanged from that represented by Figures 23B and 23C. The bandwidth of the gaps enforced here is still perceptible. However, the advantage is a very rapid measurement of the optimized frequency range, especially when gaps are enforced in multiple frequency bands at once.
[0287] Figure 24 shows another example of an audio environment. In Figure 24, environment 2409 (acoustic space) includes a user (2401) who directly speaks 2402 and an example system including smart audio devices (2403 and 2405), a speaker for audio output, and a set of microphones. The system may be configured according to an embodiment of the present disclosure. Utterances uttered by user 2401 (sometimes referred to herein as the speaker) may be recognized by elements or elements of the system in controlled time-frequency gaps.
[0288] More specifically, the elements of the system of FIG. 24 include: 2402: Direct local voice (generated by user 2401); 2403: A voice assistant device (coupled to one or more loudspeakers). Device 2403 is located closer to user 2401 than device 2405, and thus device 2403 is sometimes referred to as the "near" device and device 2405 as the "far" device; 2404: Multiple microphones in (or coupled to) nearby device 2403; 2405: Voice assistant device (coupled to one or more loudspeakers); 2406: Multiple microphones in (or coupled to) remote device 2405; 2407: Household appliances (e.g. lamps); 2408: A plurality of microphones within (or coupled to) home appliance 2407. In some examples, each of microphones 2408 may be configured to communicate with a device configured to implement a classifier, which may be at least one of devices 2403 or 2405 in some cases.
[0289] 24 may also include at least one classifier. For example, device 2403 (and / or device 2405) may include the classifier. Alternatively or additionally, the classifier may be implemented by another device that may be configured to communicate with device 2403 and / or 2405. In some examples, the classifier may be implemented by another local device (e.g., a device within environment 2409), while in other examples, the classifier may be implemented by a remote device (e.g., a server) located outside environment 2409.
[0290] In some implementations, a control system (e.g., control system 160 of FIG. 1B) may be configured to implement a classifier, such as those disclosed herein. Alternatively or additionally, control system 160 may be configured to determine an estimate of a user zone in which a user is currently located based at least in part on output from the classifier.
[0291] 25A is a flow diagram outlining one example of a method that may be performed by an apparatus such as that shown in FIG. 1B. The blocks of method 2500, as well as other methods described herein, are not necessarily performed in the order shown. Furthermore, such methods may include more or fewer blocks than those shown and / or described. In this implementation, method 2500 involves estimating a user's position within an environment.
[0292] In this example, block 2505 involves receiving an output signal from each of a plurality of microphones in the environment, where each of the plurality of microphones is present at a microphone location in the environment. According to this example, the output signal corresponds to a user's current speech measured during a controlled gap in the playback content. Block 2505 may involve, for example, a control system (such as control system 160 of FIG. 1B) receiving the output signal from each of the plurality of microphones in the environment via an interface system (such as interface system 155 of FIG. 1B).
[0293] In some examples, at least some of the microphones in the environment may provide output signals that are asynchronous with respect to output signals provided by one or more other microphones. For example, a first microphone of the plurality of microphones may sample audio data according to a first sample clock, and a second microphone of the plurality of microphones may sample audio data according to a second sample clock. In some instances, at least one of the microphones in the environment may be included within or configured to communicate with a smart audio device.
[0294] According to this example, block 2510 involves determining a plurality of current acoustic features from the output signal of each microphone. In this example, the “current acoustic features” are acoustic features derived from the “current utterance” of block 2505. In some implementations, block 2510 may involve receiving the plurality of current acoustic features from one or more other devices. For example, block 2510 may involve receiving at least some of the plurality of current acoustic features from one or more speech detectors implemented by the one or more other devices. Alternatively or additionally, in some implementations, block 2510 may involve determining the plurality of current acoustic features from the output signal.
[0295] Regardless of whether the acoustic features are determined by a single device or multiple devices, the acoustic features may be determined asynchronously. When the acoustic features are determined by multiple devices, the acoustic features are generally determined asynchronously unless the devices are configured to coordinate the process of determining the acoustic features. When the acoustic features are determined by a single device, in some implementations, the acoustic features may still be determined asynchronously because the single device may receive the output signals of each microphone at different times. In some examples, the acoustic features may be determined asynchronously because at least some of the microphones in the environment may provide output signals that are asynchronous with respect to the output signals provided by one or more other microphones.
[0296] In some examples, the acoustic features may include a speech confidence metric corresponding to speech measured during a controlled gap in the output playback signal.
[0297] Alternatively or additionally, the acoustic features may include one or more of the following: Band power in frequency bands weighted for human speech. For example, acoustic features may be based only on a specific frequency band (e.g., 400 Hz to 1.5 kHz). In this example, higher and lower frequencies may be ignored. Per-band or per-bin voice activity detector confidence in frequency bands or bins corresponding to the controlled gaps in the playback content. The acoustic features may be based at least in part on long-term noise estimates so as to ignore microphones with poor signal-to-noise ratios. Kurtosis as an indicator of speech peakiness. Kurtosis can be an indicator of smearing due to long reverberation tails.
[0298] According to this example, block 2515 involves applying a classifier to the plurality of current acoustic features. In some such examples, applying the classifier may involve applying a model trained with previously determined acoustic features derived from a plurality of previous utterances made by the user in a plurality of user zones within the environment. Various examples are provided herein.
[0299] In some examples, the user zones may include a sink area, a food preparation area, a refrigerator area, a dining area, a couch area, a television area, a sleeping area, and / or a doorway area. According to some examples, one or more of the user zones may be predetermined user zones. In some such examples, the one or more predetermined user zones may be selectable by a user during a training process.
[0300] In some implementations, applying the classifier may involve applying a Gaussian mixture model trained on previous utterances. According to some such implementations, applying the classifier may involve applying a Gaussian mixture model trained on one or more of the normalized speech confidence, normalized average received level, or maximum received level of the previous utterances. However, in alternative implementations, applying the classifier may be based on a different model, such as one of the other models disclosed herein. In some instances, the model may be trained using training data labeled with user zones. However, in some examples, applying the classifier involves applying a model trained using unlabeled training data that is not labeled with user zones.
[0301] In some instances, the previous utterance may have been or may have included a speech utterance. According to some such instances, the previous utterance and the current utterance may have been utterances of the same speech.
[0302] In this example, block 2520 involves determining an estimate of a user zone in which the user is currently located based at least in part on the output from the classifier. In some such examples, the estimate may be determined without reference to the geometric locations of multiple microphones. For example, the estimate may be determined without reference to the coordinates of individual microphones. In some examples, the estimate may be determined without estimating the geometric location of the user. However, in alternative implementations, the location estimation may involve estimating the geometric localization of one or more people and / or one or more audio devices within the audio environment, for example, with reference to a coordinate system.
[0303] Some implementations of method 2500 may involve selecting at least one speaker according to the estimated user zone. Some such implementations may involve controlling the at least one selected speaker to provide sound to the estimated user zone. Alternatively or additionally, some implementations of method 2500 may involve selecting at least one microphone according to the estimated user zone. Some such implementations may involve providing a signal output by the at least one selected microphone to a smart audio device.
[0304] FIG. 25B is a block diagram of elements of an example embodiment configured to implement a zone classifier. According to this example, a system 2530 includes multiple loudspeakers 2534 distributed throughout at least a portion of an environment (e.g., an environment such as that shown in FIG. 21 or FIG. 24). In this example, the system 2530 includes a multi-channel loudspeaker renderer 2531. According to this implementation, the output of the multi-channel loudspeaker renderer 2531 serves as both the loudspeaker drive signals (speaker feeds for driving the speakers 2534) and the echo reference. In this implementation, the echo reference is provided to an echo management subsystem 2533 via multiple loudspeaker reference channels 2532 that include at least some of the speaker feed signals output from the renderer 2531.
[0305] In this implementation, system 2530 includes multiple echo management subsystems 2533. In this example, renderer 2531, echo management subsystem 2533, wake word detector 2536, and classifier 2537 are implemented via instances of control system 160 described above with reference to FIG. 1B. According to this example, echo management subsystems 2533 are configured to implement one or more echo suppression processes and / or one or more echo cancellation processes. In this example, each of echo management subsystems 2533 provides a corresponding echo management output 2533A to one of the wake word detectors 2536. Echo management output 2533A has an attenuated echo relative to the input to the associated one of echo management subsystems 2533.
[0306] According to this implementation, the system 2530 includes N microphones 2535 (N is an integer) distributed throughout at least a portion of an audio environment (e.g., the audio environment shown in FIG. 21 or FIG. 24). The microphones may include array microphones and / or spot microphones. For example, one or more smart audio devices located within the environment may include an array of microphones. In this example, the outputs of the microphones 2535 are provided as inputs to the echo management subsystems 2533. According to this implementation, each of the echo management subsystems 2533 captures the output of an individual microphone 2535 or an individual group or subset of the microphones 2535.
[0307] In this example, system 2530 includes multiple wake word detectors 2536. According to this example, each of the wake word detectors 2536 receives audio output from one of the echo management subsystems 2533 and outputs multiple acoustic features 2536A. The acoustic features 2536A output from each echo management subsystem 2533 may include (but are not limited to) wake word confidence, wake word duration, and receive level measurements. Although three arrows representing three acoustic features 2536A are shown as being output from each echo management subsystem 2533, more or fewer acoustic features 2536A may be output in alternative implementations. Furthermore, although these three arrows are incident on classifier 2537 along more or less vertical lines, this does not indicate that classifier 2537 necessarily receives acoustic features 2536A from all wake word detectors 2536 simultaneously. As mentioned elsewhere herein, the acoustic features 2536A may, in some cases, be determined and / or provided to the classifier asynchronously.
[0308] According to this implementation, system 2530 includes a zone classifier 2537, sometimes referred to as classifier 2537. In this example, the classifier receives multiple features 2536A from multiple wake word detectors 2536 for multiple (e.g., all) microphones 2535 in the environment. According to this example, output 2538 of zone classifier 2537 corresponds to an estimate of a user zone in which the user is currently located. According to some such examples, output 2538 may correspond to one or more posterior probabilities. The estimate of the user zone in which the user is currently located may be or correspond to a maximum posterior probability according to Bayesian statistics.
[0309] Next, we describe an example implementation of a classifier, which in some instances may correspond to zone classifier 2537 of FIG. 25B. i Let (n) be the microphone signal i={1...N} at discrete time n (i.e., microphone signal x i (n) are the outputs of the N microphones 2535. N signals x i (n) produces a "clean" microphone signal e at each discrete time n. i (n), where i = {1...N}. The clean signal e, designated 2533A in FIG. 25B, is generated. i (n) are fed to wake word detectors 2536 in this example, where each wake word detector 2536 receives a vector of features w i (j), where j={1...J} is the index corresponding to the jth wake word utterance. In this example, the classifier 2537 takes as input the aggregate feature set
number
[0310] In some implementations, the zone label C for k={1…K} kThe set of may correspond to a number K of different user zones in the environment. For example, the user zones may include a couch zone, a kitchen zone, a reading chair zone, etc. Some examples may define more than one zone within a kitchen or other room. For example, a kitchen area may include a sink zone, a food preparation zone, a refrigerator zone, and a dining zone. Similarly, a living room area may include a couch zone, a TV zone, a reading chair zone, one or more doorway zones, etc. Zone labels for these zones may be selectable by the user, for example, during a training phase.
[0311] In some implementations, the classifier 2537 may use, for example, a Bayesian classifier, to determine the posterior probability p(C k |W(j)) with probability p(C k |W(j)) is the user's zone C k (for the "j"th utterance and the "k"th zone, the probability of being in zone C k and for each of the utterances) are examples of the output 2538 of the classifier 2537.
[0312] According to some examples, training data may be collected (e.g., for each user zone) by prompting the user to select or define a zone, e.g., a couch zone. The training process may involve prompting the user to make a training utterance, such as a wake word, near the selected or defined zone. In the couch zone example, the training process may involve prompting the user to make a training utterance at the center and both ends of the couch. The training process may involve prompting the user to repeat the training utterance several times at each location within the user zone. The user may then be prompted to move to another user zone and continue until all designated user zones have been covered.
[0313] FIG. 26 presents a block diagram of an example system for coordinated gap insertion. The system of FIG. 26 includes audio device 2601a, which is an instance of apparatus 150 of FIG. 1B and includes control system 160 configured to implement noise estimation subsystem 64, noise compensation gain application subsystem 62, and forced gap application subsystem 70. In this example, audio devices 2601b-2601n are also present in playback environment E. In this implementation, each of audio devices 2601b-2601n is an instance of apparatus 150 of FIG. 1B and each includes a control system configured to implement instances of noise estimation subsystem 64, noise compensation subsystem 62, and forced gap application subsystem 70.
[0314] According to this example, the system of FIG. 26 also includes a leader device 2605, which is also an instance of apparatus 150 of FIG. 1B. In some examples, the leader device 2605 may be an audio device of the playback environment, such as a smart audio device. In some such examples, the leader device 2605 may be implemented via one of the audio devices 2601a-2601n. In other examples, the leader device 2605 may be another type of device, such as what is referred to herein as a smart home hub. According to this example, the leader device 2605 includes a control system configured to receive noise estimates 2610a-2610n from the audio devices 2601a-2601n and provide urgency signals 2615a-2615n to the audio devices 2601a-2601n to control respective instances of the forced gap applicator 70. In this implementation, each instance of the forced gap applicator 70 is configured to determine whether to insert a gap, and if so, what type of gap to insert, based on the urgency signal 2615a-2615n.
[0315] According to this example, the audio devices 2601a-2601n are also configured to provide current gap data 2620a-2620n to the governing device 2605, which indicates what gaps, if any, each of the audio devices 2601a-2601n has implemented. In some examples, the current gap data 2620a-2620n may indicate the sequence of gaps the audio devices are applying and the corresponding times (e.g., start times and time durations for each or all gaps). In some implementations, the control system of the governing device 2605 may be configured to maintain a data structure indicating, for example, recent gap data, which audio devices received the most recent urgency signal, etc. In the system of FIG. 26, each instance of the mandatory gap application subsystem 70 operates in response to an urgency signal 2615a-2615n, and the governing device 2605 exercises control over mandatory gap insertion based on the need for gaps in the playback signal.
[0316] According to some examples, the urgency signals 2615a-2615n are represented by a set of urgency values [U0, U1, ... U N ], where N is a predetermined number of frequency bands (in the entire frequency range of the playback signal) into which subsystem 70 may insert forced gaps (e.g., one forced gap is inserted in each of the bands), and U i is the urgency value for the "i"th band in which subsystem 70 may insert a forced gap. The urgency values in each urgency value set (corresponding to a time) may be generated according to any disclosed embodiment for determining urgency and may indicate the urgency of insertion (by subsystem 70) of a forced gap (at that time) in the N bands.
[0317] In some implementations, the urgency signals 2615a-2615n are represented by a fixed (time-invariant) set of urgency values [U0, U1, ... U2] determined by a probability distribution defining the probability of gap insertion for each of the N frequency bands. N]. According to some examples, the probability distribution is implemented using a pseudo-random mechanism so that the results (the response of each instance of subsystem 70) are deterministic (e.g., the same) across all of the receiving audio devices 2601a-2601n. Thus, in response to such a fixed set of urgency values, subsystem 70 may be configured to insert fewer forced gaps (on average) in bands with lower urgency values (i.e., lower probability values as determined by the pseudo-random probability distribution) and more forced gaps (on average) in bands with higher urgency values (i.e., higher probability values). In some implementations, urgency signals 2615a-2615n may indicate a sequence of urgency value sets [U0, U1, ... UN], e.g., different urgency value sets for different times in the sequence. Each such different urgency value set may be determined by a different pseudo-random probability distribution for each different time.
[0318] Next, a method for determining the urgency value or a signal indicative of the urgency value (U), which may be implemented in various embodiments of the disclosed pervasive listening method, will be described.
[0319] The urgency value for a frequency band indicates the need for a gap to be forced in that band. k We present three strategies for determining U k denotes the urgency of forced gap insertion in band k, and U denotes B count represents a vector containing the urgency values for all bands in the set of frequency bands: U=[U0,U1,U2,…]
[0320] The first strategy (sometimes referred to herein as Method 1) determines a fixed urgency value. This method is the simplest and simply allows the urgency vector U to be a predetermined fixed quantity. When used with a fixed perceptual freedom metric, this can be used to implement a system that randomly inserts forced gaps over time. Some such methods do not require a time-dependent urgency value supplied by the pervasive listening application. Thus: U=[u0,u1,u2,…,u X ] where X=B count (k=1 to k=B count For each value u (for k in the range k represents a predetermined fixed urgency value for the "k" band. k Setting σ to 1.0 represents equal urgency in all frequency bands.
[0321] A second strategy (sometimes referred to herein as Method 2) determines an urgency value that depends on the time elapsed since the occurrence of the previous gap. In some implementations, the urgency gradually increases over time and returns to a low value once either a forced gap or an existing gap triggers an update in the pervasive listening results (e.g., a background noise estimate update).
[0322] Therefore, the urgency value U in each frequency band (band k) k may correspond to the duration (e.g., number of seconds) since the gap is perceived (by a pervasive listener) in band k. In some examples, the urgency value U k can be determined as follows: U k (t)=min(tt g ,U max ) where t g denotes the last gap seen for band k, and U maxrepresents a tuning parameter that limits the urgency to a certain maximum size. g Note that may be updated based on the presence of gaps originally present in the playback content. For example, in noise compensation, the current noise conditions in the playback environment may determine what is considered a gap in the output playback signal. That is, for a gap to occur, the playback signal must be quieter when the environment is quiet than when the environment is noisy. Similarly, the urgency of frequency bands typically occupied by human speech becomes more important when implementing pervasive listening methods that typically rely on the occurrence or non-occurrence of speech utterances by a user in the playback environment.
[0323] The third strategy (sometimes referred to herein as Method 3) determines an event-based urgency value. In this context, "event-based" refers to relying on some event or activity (or need for information) detected or inferred to have occurred outside or within the playback environment. The urgency determined by the pervasive listening subsystem can suddenly change with the onset of new user behavior or changes in playback environment conditions. For example, such changes may cause one or more devices configured for pervasive listening to urgently need to observe background activity in order to make decisions or quickly adjust the playback experience to new conditions, or to implement changes in the general urgency or the desired density and time between gaps in each band. Table 3 below provides some examples of contexts and scenarios and corresponding event-based changes in urgency. [Table 3]
[0324] A fourth strategy (sometimes referred to herein as Method 4) determines the urgency value using a combination of two or more of Methods 1, 2, and 3. For example, each of Methods 1, 2, and 3 may be combined into a joint strategy represented by the following type of general formulation: U k (t)=u k *min(tt g ,U max )*V k where u k represents a fixed, unitless weighting factor that controls the relative importance of each frequency band, and V k represents a scalar value that is modulated in response to changes in context or user behavior that require a rapid change in urgency, and t g and U max is defined above. In some examples, the value V k is expected to remain at a value of 1.0 under normal operation.
[0325] In some examples of multi-device contexts, the enforcement gap applicators of the smart audio devices in the audio environment may cooperate in a coordinated manner to achieve an accurate estimate of the environmental noise N. In some such implementations, the decision of where to introduce enforcement gaps in time and frequency may be made by a governing device 2605 implemented by a separate governing device (such as what is referred to elsewhere herein as a smart home hub). In some alternative implementations, the decision of where to introduce enforcement gaps in time and frequency may be made by one of the smart audio devices acting as a leader (e.g., a smart audio device acting as the governing device 2605).
[0326] In some implementations, the governing device 2605 may include a control system configured to receive the noise estimates 2610a-2610n and provide gap commands to the audio devices 2601a-2601n, which may be based at least in part on the noise estimates 2610a-2610n. In some such examples, the governing device 2605 may provide the gap commands in lieu of an urgency signal. According to some such examples, the forced gap applicator 70 need not determine whether to insert a gap, and if so, what type of gap to insert, based on the urgency signal, but instead simply operates according to the gap commands.
[0327] In some such implementations, the gap command specifies the characteristics of one or more particular gaps to be inserted (e.g., frequency range or B count , Z, t1, t2, and / or t3) and the time(s) for insertion of one or more particular gaps. For example, the gap command may indicate a sequence of gaps and corresponding time intervals, such as one of those shown in FIGS. 23B-23E and described above. In some examples, the gap command may indicate a data structure from which the receiving audio device can access the characteristics of the sequence of gaps and corresponding time intervals to be inserted. The data structure may, for example, have been previously provided to the receiving audio device. In some such examples, the governing device 2605 may include a control system configured to perform an urgency calculation to determine when to send a gap command and what type of gap command to send.
[0328] According to some examples, the urgency signal may be estimated, at least in part, by noise estimation elements 64 of one or more of audio devices 2601a-2601n and transmitted to governing device 2605. The decision to govern a forced gap to a particular frequency domain and time location may, in some examples, be determined, at least in part, by an aggregation of these urgency signals from one or more of audio devices 2601a-2601n. For example, a disclosed algorithm making an urgency-informed selection may instead use a maximum urgency calculated across the urgency signals of multiple audio devices, e.g., Urgency=maximum(UrgencyA, UrgencyB, UrgencyC, ...), where UrgencyA / B / C are understood to be the urgency signals of three separate exemplary devices implementing noise compensation.
[0329] Noise compensation systems (e.g., those of FIG. 26 ) can function with weak or nonexistent echo cancellation (e.g., when implemented as described in U.S. Provisional Patent Application No. 62 / 663,302, incorporated herein by reference), but can suffer from content-dependent response times, particularly for music, TV, and movie content. The time it takes for a noise compensation system to respond to changes in the background noise profile in the playback environment can be critical to the user experience and, in some cases, may be more important than the accuracy of the actual noise estimate. When the playback content provides little or no gap in knowledge of one edge of the background noise, the noise estimate may remain fixed even as noise conditions change. While interpolating and imputing missing values in the noise estimate spectrum is typically useful, it is still possible for large regions of the noise estimate spectrum to lock up and become stale.
[0330] Some embodiments of the system of FIG. 26 may be operable to provide forced gaps (in the playback signal) that occur frequently enough (e.g., in each frequency band of interest at the output of forced gap applicator 70) so that the background noise estimate (by noise estimator 64) can be updated frequently enough to respond to typical changes in the profile of background noise N in playback environment E. In some examples, subsystem 70 may be configured to introduce forced gaps in a compensated audio playback signal (having K channels, where K is a positive integer) output from noise compensation subsystem 62. Here, noise estimator 64 may be configured to search for gaps (including the forced gaps inserted by subsystem 70) in each channel of the compensated audio playback signal and generate noise estimates for the frequency bands (and time intervals) in which the gaps occur. In this example, noise estimator 64 of audio device 2601a is configured to provide noise estimate 2610a to noise compensation subsystem 62. According to some examples, noise estimator 64 of audio device 2601 a may also be configured to use the resulting information regarding the detected gaps to generate (and provide to) governing device 2605 an estimated urgency signal, the urgency value tracking the urgency for inserting forced gaps in frequency bands of the compensated audio playback signal.
[0331] In this example, noise estimator 64 is configured to accept both microphone feed Mic (output of microphone M in playback environment E) and a reference compensated audio playback signal (input to speaker system S in playback environment E). According to this example, the noise estimate generated in subsystem 64 is provided to noise compensation subsystem 62, which applies compensation gains to input playback signal 23 (from content source 22) to level each frequency band thereof to a desired playback level. In this example, the noise-compensated audio playback signal (output from subsystem 62) and a per-band urgency metric (indicated by an urgency signal output from governing device 2605) are provided to forced gap applicator 70, which enforces gaps in the compensated playback signal (preferably according to an optimization process). Speaker feeds (output from forced gap applicator 70), each indicating the content of a different channel of the noise-compensated playback signal, are provided to each speaker of speaker system S.
[0332] While some implementations of the system of Figure 26 may perform echo cancellation as part of the noise estimation they perform, other implementations of the system of Figure 26 do not perform echo cancellation, and therefore elements for implementing echo cancellation are not specifically shown in Figure 26.
[0333] While FIG. 26 does not show the transformation of the signal from the time domain to the frequency domain (and / or from the frequency domain to the time domain), the application of noise compensation gain (in subsystem 62), the analysis of content for gap enforcement (in governing device 2605, noise estimator 64, and / or forced gap applicator 70), and the insertion of forced gaps (by forced gap applicator 70) may conveniently be implemented in the same transform domain, with the resulting output audio being resynthesized into time-domain pulse code modulation (PCM) audio before further encoding for playback or transmission. According to some examples, each participating device coordinates the enforcement of such gaps using methods described elsewhere herein. In some such examples, the introduced gaps may be identical. In some examples, the introduced gaps may be synchronized.
[0334] By using a forced gap applicator 70 present on each participating device to insert gaps, the number of gaps in each channel of the compensated playback signal (output from the noise compensation subsystem 62 of the system of FIG. 26) can be increased (compared to the number of gaps that would occur without the use of the forced gap applicator 70), thereby significantly reducing the requirements for any echo canceller implemented by the system of FIG. 26, and in some cases even eliminating the need for echo cancellation altogether.
[0335] In some disclosed implementations, simple post-processing circuitry, such as time-domain peak limiting or speaker protection, can be implemented between the forced gap applicator 70 and the speaker system S. However, post-processing capable of boosting and compressing the speaker feeds can counteract or degrade the quality of the forced gaps inserted by the forced gap applicator, and therefore these types of post-processing are preferably implemented at some point in the signal processing path before the forced gap applicator 70.
[0336] 27A and 27B show system block diagrams illustrating example elements of a governing device and elements of governed audio devices, according to some disclosed implementations. As with other figures provided herein, the types and number of elements shown in FIGS. 27A and 27B are provided merely as examples. Other implementations may include more, fewer, different types, and / or different numbers of elements. In this example, governed audio devices 2720a-2720n and governing device 2701 of FIGS. 27A and 27B are instances of apparatus 150 described above with reference to FIG. 1B.
[0337] According to this implementation, each of the governed audio devices 2720a-2720n includes the following elements: 2731: an instance of loudspeaker system 110 of FIG. 1B including one or more loudspeakers; 2732: an instance of microphone system 111 of FIG. 1B including one or more microphones; 2711: An audio playback signal output by a rendering module 2721, which in this example is an instance of rendering module 210A of Figure 2. According to this example, rendering module 2721 is controlled according to instructions from governing module 2702 and may also receive information and / or instructions from user zone classifier 2705 and / or rendering configuration module 2707; 2712: a noise compensated audio playback signal output by noise compensation module 2721, which in this example is an instance of noise compensation subsystem 62 of FIG. 26; 2713: A noise-compensated audio playback signal including one or more gaps output by an acoustic gap puncher 2722, which in this example is an instance of the forced gap applicator 70 of Figure 26. In this example, the acoustic gap puncher 2722 is controlled according to instructions from the leadership module 2702; 2714: modified audio playback signal output by calibration signal injector 2723, which in this example is an instance of calibration signal injector 211A of FIG. 2; 2715: calibration signal output by calibration signal generator 2725, which in this example is an instance of calibration signal generator 212A of FIG. 2; 2716: A calibration signal replica corresponding to a calibration signal generated by another audio device in the audio environment (in this example, by one or more of audio devices 2720b-2720n). Calibration signal replica 2716 may be, for example, an instance of calibration signal replica 204A described above with reference to FIG. 2. In some examples, calibration signal replica 2716 may be received from the leader device 2701 (e.g., via a wireless communication protocol such as Wi-Fi or Bluetooth); 2717: Control information associated with and / or used by one or more of the audio devices in the audio environment. In this example, the control information 2717 is provided by the governing device 2701 (e.g., by the governing module 2702), described below with reference to FIG. 27B. The control information 2717 may include, for example, an instance of the calibration information 205A described above with reference to FIG. 2, or an instance of a calibration signal parameter disclosed elsewhere herein. The control information 2717 may include parameters used by the control system 160n to generate a calibration signal, modulate a calibration signal, demodulate a calibration signal, etc. The control information 2717, in some examples, may include one or more DSSS spreading code parameters and one or more DSSS carrier parameters. The control information 2717, in some examples, may include information for controlling the rendering module 2721, the noise compensation module 2711, the acoustic gap puncher 2712, and / or the baseband processor 2729; 2718: microphone signal received by microphone 2732; 2719: Demodulated coherent baseband signal, which may be an instance of the demodulated coherent baseband signals 208 and 208A described above with reference to Figures 2-4 and 17; 2721: A rendering module configured to render an audio signal of a content stream, such as audio data for music, movies, and television programs, to generate an audio playback signal; 2723: a calibration signal injector configured to insert the calibration signal 2715a modulated by the calibration signal modulator 2724 (or, in some cases where the calibration signal does not require modulation, the calibration signal 2715 generated by the calibration signal generator 2725) into the audio playback signal generated by the rendering module 2721 (which in this example has been modified by the noise compensation module 2730 and the acoustic gap puncher 2722) to generate the modified audio playback signal 2714. The insertion process may, for example, be a mixing process in which the calibration signal 2715 or 2715a is mixed with the audio playback signal generated by the rendering module 210A (which in this example has been modified by the noise compensation module 2730 and the acoustic gap puncher 2722) to generate the modified audio playback signal 2714; 2724: optional calibration signal modulator configured to modulate the calibration signal 2715 generated by the calibration signal generator 2725 to generate a modulated calibration signal 2715a; 2725: A calibration signal generator configured to generate the calibration signal 2715 and, in this example, provide the calibration signal 2715 to the calibration signal modulator 2724 and the baseband processor 2729. In some examples, the calibration signal generator 2725 may be an instance of the calibration signal generator 212A described above with reference to FIG. 2. According to some examples, the calibration signal generator 2725 may include a spreading code generator and a carrier generator, for example, as described above with reference to FIG. 17. In this example, the calibration signal generator 2725 provides the calibration signal replica 2715 to the baseband processor and the calibration signal demodulator 2726; 2726: A calibration signal demodulator configured to demodulate the microphone signal 2718 received by the microphone 2732. In some examples, the calibration signal demodulator 2726 may be an instance of the calibration signal demodulator 212A described above with reference to FIG. 2. In this example, the calibration signal demodulator 2726 outputs a demodulated coherent baseband signal 2719. Demodulation of the microphone signal 2718 may be performed using standard correlation techniques, including, for example, an integrate-and-dump matched filtering correlator bank. Several detailed examples are provided herein. To improve the performance of these demodulation techniques, in some implementations, the microphone signal 2718 may be filtered before demodulation to remove undesired content / phenomena. According to some implementations, the demodulated coherent baseband signal 2719 may be filtered before or after being provided to the baseband processor 2729. The signal-to-noise ratio (SNR) generally improves as the integration time increases (e.g., as the length of the spreading code used to generate the calibration signal increases); 2729: A baseband processor configured for baseband processing of the demodulated coherent baseband signal 2719. In some examples, the baseband processor 2729 may be configured to implement techniques such as incoherent averaging to improve the SNR by reducing the variance of the squared waveform to generate the delayed waveform. Several detailed examples are provided herein. In this example, the baseband processor 218A is configured to output one or more estimated acoustic scene metrics 2733; 2730: a noise compensation module configured to compensate for noise in the audio environment. In this example, the noise compensation module 2730 compensates for noise in the audio playback signal 2711 output by the rendering module 2721 based at least in part on control information 2717 from the governing module 2702. In some implementations, the noise compensation module 2730 may be configured to compensate for noise in the audio playback signal 2711 based at least in part on one or more acoustic scene metrics 2733 (e.g., noise information) provided by the baseband processor 2729; 2733n: One or more observations derived by audio device 2720n, for example, from a calibration signal extracted from a microphone signal (e.g., from demodulated coherent baseband signal 2719) and / or from wake word information 2734 provided by wake word detector 2727. These observations are also referred to herein as acoustic scene metrics. Acoustic scene metrics 2733 may include or be wake word metrics, data corresponding to time of flight, time of arrival, range, audio device audibility, audio device impulse response, angle between audio devices, audio device position, audio environment noise, and / or signal-to-noise ratio. In this example, governed audio devices 2720a-2720n have determined acoustic scene metrics 2733a-2733n, respectively, and provide the acoustic scene metrics 2733a-2733n to governing device 2701.
[0338] According to this implementation, the leader device 2701 includes the following elements: 2702: A coordinating module configured to control various functions of the coordinated audio devices 2720a-2720n, including, but not limited to, gap insertion and calibration signal generation in this example. The coordinating module 2702, in some implementations, may provide one or more of the various functions of the coordinated device disclosed herein. Thus, the coordinating module 2702 may provide information for controlling one or more aspects of audio processing and / or audio device playback. For example, the coordinating module 2702 may provide calibration signal parameters to the calibration signal generators 2725 (and, in this example, the modulators 2724 and demodulators 2726) of the coordinated audio devices 2720a-2720n. The coordinating module 2702 may provide gap insertion information to the acoustic gap punchers 2722 of the coordinated audio devices 2720a-2720n. The coordinating module 2702 may provide instructions for coordinating gap insertion and calibration signal generation. The governing module 2702 (and in some examples, other modules of the governing device 2701, e.g., the user zone classifier 2705 and the rendering configuration generator 2707 in this example) may provide instructions to control the rendering module 2721; 2703: A geometric proximity estimator configured to estimate the current position, and in some instances, the current orientation, of an audio device within the audio environment. In some instances, the geometric proximity estimator 2703 may be configured to estimate the current position (and in some instances, the current orientation) of one or more people within the audio environment. Some examples of the functionality of the geometric proximity estimator are described below with reference to FIG. 41 et seq.; 2704: Audio device audibility estimator that may be configured to estimate the audibility of one or more loudspeakers in or near an audio environment at any location, e.g., at a listener's current estimated location. Some examples of the functionality of the audio device audibility estimator are described below with reference to Figure 31 et seq. (see, e.g., Figure 32 and corresponding discussion); 2705: a user zone classifier configured to estimate a zone of the audio environment in which a person is currently located (e.g., couch zone, kitchen table zone, refrigerator zone, reading chair zone, etc.). In some examples, user zone classifier 2705 may be an instance of zone classifier 2537, the functionality of which is described above with reference to Figures 25A and 25B; 2706: A noise audibility estimator configured to estimate noise audibility at any position, the audibility at a listener's current estimated position in the audio environment. Some examples of the functionality of the audio device audibility estimator are described below with reference to FIG. 31 et seq. (see, e.g., FIG. 33 and FIG. 34 and corresponding descriptions). The noise audibility estimator 2706 may, in some examples, estimate noise audibility by interpolating aggregated noise data 2740 from the aggregator 2708. The aggregated noise data 2740 may be obtained, for example, from multiple audio devices in the audio environment (e.g., by multiple baseband processors 2729 and / or other modules implemented by the control systems of the audio devices), for example, by "listening through" gaps inserted in the played audio data to assess the noise conditions in the audio environment, as described above with reference to FIG. 21 et seq.; 2707: A rendering configuration generator configured to generate a rendering configuration in response to the relative positions (and, in this example, relative audibility) of audio devices and one or more listeners in the audio environment. Rendering configuration generator 2707 may provide, for example, functionality as described below with reference to FIG. 51 et seq.; 2708: an aggregator configured to aggregate acoustic scene metrics 2733a-2733n received from the governed audio devices 2701a-2720n and provide aggregated acoustic scene metrics (in this example, aggregated acoustic scene metrics 2735-2740) to the acoustic scene metric processing module 2728 and other modules of the governing device 2720. Because acoustic scene metric estimates from the baseband processor modules of the governed audio devices 2720a-2720n generally arrive asynchronously, the aggregator 2708 is configured to collect the acoustic scene metric data over time, store the acoustic scene metric data in a memory (e.g., a buffer), and pass it to subsequent processing blocks at the appropriate time (e.g., after acoustic scene metric data has been received from all governed audio devices). In this example, the aggregator 2708 is configured to provide aggregated audibility data 2735 to the governing module 2702 and the audio device audibility estimator 2704. In this implementation, the aggregator 2708 is configured to provide aggregated noise data 2740 to the governing module 2702 and the noise audibility estimator 2706. According to this implementation, the aggregator 2708 provides aggregated direction of arrival (DOA) data 2736, aggregated time of arrival (TOA) data 2737, and aggregated impulse response (IR) data 2738 to the governing module 2702 and the geometric proximity estimator 2703. In this example, the aggregator 2708 provides aggregated wake word metrics 2739 to the governing module 2702 and the user zone classifier 2705; 2728: an acoustic scene metric processing module configured to receive and apply aggregated acoustic scene metrics 2735-2739. According to this example, acoustic scene metric processing module 2728 is a component of governing module 2702, although in alternative examples, acoustic scene metric processing module 2728 may not be a component of governing module 2702. In this example, acoustic scene metric processing module 2728 is configured to generate information and / or commands based at least in part on at least one of the aggregated acoustic scene metrics 2735-2739 and / or at least one audio device characteristic. The audio device characteristic may be one or more characteristics of one or more of the governed audio devices 2720a-2720n. The audio device characteristic may be stored in a memory of or accessible to control system 160 of governing device 2701, for example.
[0339] In some implementations, the leader device 2701 may be implemented in an audio device, such as a smart audio device, and may include one or more microphones and one or more loudspeakers.
[0340] Cloud Processing In some implementations, the coordinated audio devices 2720a-2720n include primarily real-time processing blocks that execute locally due to high data bandwidth and low processing latency requirements. However, in some examples, the baseband processor 2729 may reside in the cloud (e.g., implemented via one or more servers) because the output of the baseband processor 2729 may be calculated asynchronously in some examples. According to some implementations, all of the blocks of the coordinated device 2701 may reside in the cloud. In some alternative implementations, blocks 2702, 2703, 2708, and 2705 may be implemented on a local device (e.g., a device in the same audio environment as the coordinated audio devices 2720a-2720n) because it is preferable for these blocks to operate in real time or near real time. However, in some such implementations, blocks 2703, 2704, and 2707 may operate via a cloud service.
[0341] FIG. 28 is a flow diagram outlining another example of the disclosed audio device governance method. The blocks of method 2800, as well as other methods described herein, are not necessarily performed in the order shown. Moreover, such methods may include more or fewer blocks than those shown and / or described. Method 2800 may be performed by a governing device, such as governing device 2701 described above with reference to FIG. 27B. Method 2800 involves controlling governed audio devices, such as some or all of governed audio devices 2720a-2720n described above with reference to FIG. 27A.
[0342] According to this example, block 2805 involves causing a first audio device of the audio environment, by a control system, to generate a first calibration signal. For example, a control system of a governing device, such as governing device 2701, may be configured to cause a first governed audio device of the audio environment (e.g., governed audio device 2720a) to generate a first calibration signal in block 2805.
[0343] In this example, block 2810 involves causing the control system to insert a first calibration signal into a first audio playback signal corresponding to a first content stream to generate a first modified audio playback signal for a first audio device. For example, the governing device 2701 may be configured to cause a governed audio device 2720a to insert a first calibration signal into a first audio playback signal corresponding to the first content stream to generate a first modified audio playback signal for the governed audio device 2720a.
[0344] According to this example, block 2815 involves causing the control system to cause the first audio device to play the first modified audio playback signal to generate the first audio device playback sound. For example, the governing device 2701 may be configured to cause the governed audio device 2720a to play the first modified audio playback signal on the loudspeaker 2731 to generate the first governed audio device playback sound.
[0345] In this example, block 2820 involves causing a second audio device in the audio environment, by the control system, to generate a second calibration signal. For example, the governing device 2701 may be configured to cause the governed audio device 2720b to generate a second calibration signal.
[0346] According to this example, block 2825 includes causing the control system to insert the second calibration signal into the second content stream to generate a second modified audio playback signal for the second audio device. For example, the governing device 2701 may be configured to cause the governed audio device 2720b to insert the second calibration signal into the second content stream to generate a second modified audio playback signal for the governed audio device 2720b.
[0347] In this example, block 2830 includes causing the control system to cause the second audio device to play the second modified audio playback signal to generate the second audio device playback sound. For example, the governing device 2701 may be configured to cause the governed audio device 2720b to play the second modified audio playback signal on the loudspeaker 2731 to generate the second governed audio device playback sound.
[0348] According to this example, block 2835 involves causing, by the control system, at least one microphone in the audio environment to detect at least the first audio device playback sound and the second audio device playback sound and generate microphone signals corresponding to at least the first audio device playback sound and the second audio device playback sound. In some examples, the microphone may be a microphone of the coordinating device. In other examples, the microphone may be a microphone of a coordinated audio device. For example, the coordinating device 2701 may be configured to cause one or more of the coordinated audio devices 2720a-2720n to detect at least the first coordinated audio device playback sound and the second coordinated audio device playback sound using at least one microphone and generate microphone signals corresponding to at least the first coordinated audio device playback sound and the second coordinated audio device playback sound.
[0349] In this example, block 2840 involves causing the control system to extract the first and second calibration signals from the microphone signals. For example, the governing device 2701 may be configured to cause one or more of the governed audio devices 2720a-2720n to extract the first and second calibration signals from the microphone signals.
[0350] According to this example, block 2845 involves causing the control system to estimate at least one acoustic scene metric based at least in part on the first calibration signal and the second calibration signal. For example, the governing device 2701 may be configured to cause one or more of the governed audio devices 2720a-2720n to estimate at least one acoustic scene metric based at least in part on the first calibration signal and the second calibration signal. Alternatively or additionally, in some examples, the governing device 2701 may be configured to estimate the acoustic scene metric based at least in part on the first calibration signal and the second calibration signal.
[0351] The particular acoustic scene metrics estimated in method 2800 may vary according to the particular implementation. In some examples, the acoustic scene metrics may include one or more of time of flight, time of arrival, direction of arrival, range, audio device audibility, audio device impulse response, angle between audio devices, audio device position, audio environment noise, or signal-to-noise ratio.
[0352] In some examples, the first calibration signal may correspond to a first inaudible component of sound reproduced by a first audio device, and the second calibration signal may correspond to a second inaudible component of sound reproduced by a second audio device.
[0353] In some cases, the first calibration signal may be or include a first DSSS signal, and the second calibration signal may be or include a second DSSS signal, but the first and second calibration signals may be any suitable type of calibration signal, including but not limited to the examples disclosed herein.
[0354] According to some examples, a first content stream component of sound reproduced by a first coordinated audio device may cause perceptual masking of a first calibration signal component of sound reproduced by the first coordinated audio device, and a second content stream component of sound reproduced by a second coordinated audio device may cause perceptual masking of a second calibration signal component of sound reproduced by the second coordinated audio device.
[0355] In some implementations, method 2800 may involve causing a control system to insert a first gap in a first frequency range of the first audio playback signal or the first modified audio playback signal during a first time interval of the first content stream, such that the first modified audio playback signal and the first audio device playback sound include the first gap. The first gap may correspond to an attenuation of the first audio playback signal in the first frequency range. For example, the governing device 2701 may be configured to cause the governed audio device 2720a to insert a first gap in a first frequency range of the first audio playback signal or the first modified audio playback signal during the first time interval.
[0356] According to some implementations, method 2800 may involve causing a control system to insert the first gap in the first frequency range of the second audio playback signal or the second modified audio playback signal during the first time interval, such that the second modified audio playback signal and the second audio device playback sound include the first gap. For example, governing device 2701 may be configured to cause governed audio device 2720b to insert a first gap in the first frequency range of the second audio playback signal or the second modified audio playback signal during the first time interval.
[0357] In some implementations, method 2800 may involve causing a control system to extract audio data from the microphone signal within at least the first frequency range to generate extracted audio data. For example, the governing device 2701 may cause one or more of the governed audio devices 2720a-2720n to extract audio data from the microphone signal within at least the first frequency range to generate extracted audio data.
[0358] According to some implementations, method 2800 may involve causing a control system to estimate at least one acoustic scene metric based at least in part on the extracted audio data. For example, the leader device 2701 may cause one or more of the leader audio devices 2720a-2720n to estimate at least one acoustic scene metric based at least in part on the extracted audio data. Alternatively or additionally, in some examples, the leader device 2701 may be configured to estimate the acoustic scene metric based at least in part on the extracted audio data.
[0359] Method 2800 may involve controlling both gap insertion and calibration signal generation. In some examples, method 2800 may involve controlling gap insertion and / or calibration signal generation such that the perceived level of the reproduced audio content at a user's location is maintained, possibly under varying noise conditions (e.g., varying noise spectra). According to some examples, method 2800 may involve controlling calibration signal generation such that the signal-to-noise ratio of the calibration signal is maximized. Method 2800 may involve controlling calibration signal generation to ensure that the calibration signal is inaudible to the user, even under varying audio content and noise conditions.
[0360] In some examples, method 2800 may involve controlling gap insertion to empty time-frequency tiles such that neither content nor calibration signal is present during the inserted gap, thereby allowing background noise to be estimated. Thus, in some examples, method 2800 may involve controlling gap insertion and calibration signal generation such that the calibration signal does not correspond to either a gap time interval or a gap frequency range. For example, governing device 2701 may be configured to control gap insertion and calibration signal generation such that the calibration signal does not correspond to either a gap time interval or a gap frequency range.
[0361] According to some examples, method 2800 may involve controlling gap insertion and calibration signal generation based at least in part on the time since noise was estimated in at least one frequency band. For example, governing device 2701 may be configured to control gap insertion and calibration signal generation based at least in part on the time since noise was estimated in at least one frequency band.
[0362] In some examples, method 2800 may involve controlling gap insertion and calibration signal generation based at least in part on a signal-to-noise ratio of a calibration signal of at least one audio device in at least one frequency band. For example, governing device 2701 may be configured to control gap insertion and calibration signal generation based at least in part on a signal-to-noise ratio of a calibration signal of at least one governed audio device in at least one frequency band.
[0363] According to some implementations, method 2800 may involve causing a target audio device to play an unmodified audio playback signal of the target device content stream to generate a target audio device playback sound. In some such examples, method 2800 may involve estimating at least one of a target audio device audibility or a target audio device location based at least in part on the extracted audio data. In some such implementations, the unmodified audio playback signal does not include the first gap. In some such examples, the microphone signal also corresponds to the target audio device playback sound. According to some such examples, the unmodified audio playback signal does not include an inserted gap in any frequency range.
[0364] For example, the governing device 2701 may be configured to cause a target governed audio device among the governed audio devices 2720a-2720n to play an unmodified audio playback signal of the target device content stream to generate a target governed audio device playback sound. In one example, if the target audio device is governed audio device 2720a, the governing device 2701 may cause governed audio device 2720a to play an unmodified audio playback signal of the target device content stream to generate a target governed audio device playback sound. The governing device 2701 may also be configured to cause at least one of the other governed audio devices (in the examples described above, one or more of governed audio devices 2720b-2720n) to estimate at least one of the audibility or location of the target governed audio device based at least in part on the extracted audio data. Alternatively or additionally, in some examples, the governing device 2701 may be configured to estimate the audibility of a target governed audio device and / or the location of a target governed audio device based at least in part on the extracted audio data.
[0365] In some examples, the method 2800 may involve controlling one or more aspects of audio device playback based at least in part on acoustic scene metrics. For example, the governing device 2701 may be configured to control the rendering modules 2721 of one or more of the governed audio devices 2720b-2720n based at least in part on the acoustic scene metrics. In some implementations, the governing device 2701 may be configured to control the noise compensation modules 2730 of one or more of the governed audio devices 2720b-2720n based at least in part on the acoustic scene metrics.
[0366] According to some implementations, method 2800 may involve causing third through Nth audio devices of the audio environment to generate third through Nth calibration signals, and causing the control system to insert the third through Nth calibration signals into the third through Nth content streams to generate third through Nth modified audio playback signals for the third through Nth audio devices. In some examples, method 2800 may involve causing the control system to cause the third through Nth audio devices to play corresponding instances of the third through Nth modified audio playback signals to generate third through Nth instances of audio device playback sounds. For example, the governing device 2701 may be configured to cause the governed audio devices 2720c-2720n to generate third through Nth calibration signals and insert the third through Nth calibration signals into third through Nth content streams to generate third through Nth modified audio playback signals for the governed audio devices 2720c-2720n. The governing device 2701 may be configured to cause the governed audio devices 2720c-2720n to play corresponding instances of the third through Nth modified audio playback signals to generate third through Nth instances of audio device playback sound.
[0367] In some examples, method 2800 may involve causing, by the control system, at least one microphone of each of first through Nth audio devices to detect first through Nth instances of audio device reproduced sound and generate microphone signals corresponding to the first through Nth instances of audio device reproduced sound. In some instances, the first through Nth instances of audio device reproduced sound may include a first audio device reproduced sound, a second audio device reproduced sound, and third through Nth instances of audio device reproduced sound. According to some examples, method 2800 may involve causing the control system to extract first through Nth calibration signals from the microphone signals. Acoustic scene metrics may be estimated based at least in part on the first through Nth calibration signals.
[0368] For example, the leader device 2701 may be configured to cause at least one microphone of some or all of the leader audio devices 2720a-2720n to detect first through Nth instances of audio device playback sound and generate microphone signals corresponding to the first through Nth instances of audio device playback sound. The leader device 2701 may be configured to cause some or all of the leader audio devices 2720a-2720n to extract first through Nth calibration signals from the microphone signals. Some or all of the leader audio devices 2720a-2720n may be configured to estimate acoustic scene metrics based at least in part on the first through Nth calibration signals. Alternatively or additionally, the leader device 2701 may be configured to estimate acoustic scene metrics based at least in part on the first through Nth calibration signals.
[0369] According to some implementations, method 2800 may involve determining one or more calibration signal parameters for multiple audio devices in an audio environment. The one or more calibration signal parameters may be usable to generate a calibration signal. Method 2800 may involve providing the one or more calibration signal parameters to one or more governed audio devices in the audio environment. For example, governing device 2701 (and in some examples, governing module 2702 of governing device 2701) may be configured to determine one or more calibration signal parameters for one or more of governed audio devices 2720a-2720n and provide the one or more calibration signal parameters to the governed audio devices.
[0370] In some examples, determining the one or more calibration signal parameters may involve scheduling a time slot for each audio device of a plurality of audio devices to play the modified audio playback signal, and in some instances, a first time slot for a first audio device may be different from a second time slot for a second audio device.
[0371] According to some implementations, determining the one or more calibration signal parameters may involve determining a frequency band for each audio device of a plurality of audio devices to reproduce a modified audio playback signal. In some examples, a first frequency band for a first audio device may be different from a second frequency band for a second audio device.
[0372] In some examples, determining the one or more calibration signal parameters may involve determining a DSSS spreading code for each audio device of a plurality of audio devices. According to some examples, a first spreading code for a first audio device may be different from a second spreading code for a second audio device. According to some implementations, method 2800 may involve determining at least one spreading code length based at least in part on the audibility of a corresponding audio device.
[0373] In some implementations, determining the one or more calibration signal parameters may involve applying an acoustic model based at least in part on the interaudibility of each of a plurality of audio devices in the audio environment.
[0374] In some examples, the method 2800 may involve causing each of multiple audio devices in an audio environment to simultaneously play the modified audio playback signal.
[0375] According to some implementations, at least a portion of the first audio playback signal, at least a portion of the second audio playback signal, or at least a portion of each of the first audio playback signal and the second audio playback signal may correspond to silence.
[0376] FIG. 29 is a flow diagram outlining another example of the disclosed audio device governance method. The blocks of method 2900, as with other methods described herein, are not necessarily performed in the order shown. Moreover, such methods may include more or fewer blocks than those shown and / or described. Method 2900 may be performed by a governing device, such as governing device 2701 described above with reference to FIG. 27B. Method 2900 involves controlling governed audio devices, such as some or all of governed audio devices 2720a-2720n described above with reference to FIG. 27A.
[0377] The table below defines the notation used in Figure 29 and the following description. [Table 4]
[0378] In this example, Figure 29 shows blocks of a method for allocation of spectral band k in time block l. According to this example, the blocks shown in Figure 29 are repeated for each spectral band and each time block. The length of the time block may vary according to the particular implementation, but may be, for example, on the order of a few seconds (e.g., in the range of 1 to 5 seconds) or on the order of hundreds of milliseconds. The spectrum occupied by a single frequency band may also vary according to the particular implementation. In some implementations, the spectrum occupied by a single band is based on a perceptual interval, such as a Mel band or a critical band.
[0379] As used herein, the term "time-frequency tile" refers to a single block of time in a single frequency band. At any given time, the time-frequency tile may be occupied by a combination of program content (e.g., movie audio content, music, etc.) and one or more calibration signals. If only background noise needs to be sampled, neither program content nor calibration signals should be present. The corresponding time-frequency tile is referred to herein as a "gap."
[0380] The left column of Figure 29 (blocks 2902-2908) involves estimating background noise in the audio environment when neither content nor calibration signals are present in a time-frequency tile (i.e., when that time-frequency tile corresponds to a gap). This is a simplified example of the disciplined gap method as described above with reference to Figures 21 et seq., with additional logic to handle calibration sequences that may potentially occupy the same band.
[0381] In this example, the process for spectral band k in time block l begins in block 2901. Block 2902 involves determining whether the previous block (block l-1) had a gap in spectral band k. If so, this time-frequency tile corresponds only to background noise, which can be estimated in block 2903.
[0382] In this example, the noise is assumed to be quasi-stationary, so that the noise is N The noise needs to be sampled at regular intervals defined by T since the last noise measurement. N This involves determining whether a certain period has elapsed.
[0383] At block 2904, T N If it is determined that the calibration signal has elapsed, the process continues to block 2905, which involves determining whether the calibration signal in the current time-frequency tile is complete. Block 2905 is desirable because, in some implementations, the calibration signal may occupy more than one time block, and it is necessary (or at least desirable) to wait until the calibration signal in the current time-frequency tile is complete before a gap is inserted. In this example, if it is determined in block 2905 that the calibration signal is incomplete, the method proceeds to block 2906, which involves flagging the current time-frequency tile as needing a noise estimate in a future block.
[0384] In this example, if the calibration signal is determined to be complete in block 2905, the method proceeds to block 2907, which in this example G The purpose of this method is to determine whether there are any muted (gapped) frequency bands within the minimum spectral gap interval, denoted as K, so as not to produce perceptible artifacts in the reproduced audio data. GIt should be noted that block 2907 does not mute (insert gaps into) any frequency bands within the minimum spectral gap interval. If block 2907 determines that there is a gapped frequency band within the minimum spectral gap interval, the process proceeds to block 2906, where the band is flagged as requiring future noise estimation. However, if block 2907 determines that there is no gapped frequency band within the minimum spectral gap interval, the process proceeds to block 2908, which involves inserting gaps into that band by all governed audio devices. In this example, block 2908 also involves sampling the noise in the current time-frequency tile.
[0385] The right column of FIG. 29 (blocks 2909-2917) involves processing any calibration signals (also referred to herein as calibration sequences) that may have been performed in the previous time block. In some examples, each time-frequency tile may include multiple orthogonal calibration signals (such as the DSSS sequences described herein), e.g., a set of calibration signals inserted into / mixed with audio content and played by each of multiple coordinated audio devices. Thus, in this example, block 2909 involves iterating through all calibration sequences present in the current time-frequency tile to determine whether all calibration sequences have been serviced. If not, the next calibration sequence is serviced, beginning with block 2910.
[0386] Block 2911 involves determining whether the calibration sequence is complete. In some examples, the calibration sequence may span multiple time blocks, and thus a calibration sequence that started before the current time block is not necessarily complete at the time of the current time block. If the calibration sequence is determined to be complete in block 2911, the process continues to block 2912.
[0387] In this example, block 2912 involves determining whether the currently evaluated calibration sequence was successfully demodulated. Block 2912 may be based, for example, on information obtained from one or more controlled audio devices attempting to demodulate the currently evaluated calibration sequence. Demodulation failure may occur due to one or more of the following: 1. High level of background noise; 2. High-level program content; 3. High level calibration signals from nearby devices (especially the near-far problem discussed elsewhere herein); 4. Device asynchrony.
[0388] If block 2912 determines that the calibration sequence was successfully demodulated, the process proceeds to block 2913. According to this example, block 2913 involves estimating one or more acoustic scene metrics, such as DOA, TOA, and / or audibility in the current frequency band. Block 2913 may be performed by one or more coordinated devices and / or by a coordinate device.
[0389] In this example, if it is determined in block 2912 that the calibration sequence was not successfully demodulated, the process continues directly to block 2914. According to this example, block 2914 involves monitoring the demodulated calibration signal and updating the calibration signal parameters as necessary to ensure that all governed devices can hear each other well enough (have sufficiently high interaudibility). The robustness of the calibration signal parameters is determined by the parameter ζ for the i-th device in the k-th band. i,k In one example where the calibration signal is a DSSS signal, the robustness may include modifying parameters, for example, by doing one or more of the following: 1. Increase the amplitude of the calibration signal; 2. Reduce the chipping rate of the calibration signal; 3. Increase the coherent integration time; 4. Increase the incoherent integration time; and / or 5. Reduce the number of concurrent signals in the same time-frequency tile.
[0390] Calibration parameters 2 and 3 may lead to calibration sequences occupying an increased number of time blocks.
[0391] According to this example, block 2915 involves determining whether a calibration parameter has reached one or more limits. For example, block 2915 may involve determining whether the amplitude of the calibration signal has reached a limit beyond which the calibration signal will sound louder than the reproduced audio content. In some examples, block 2915 may involve determining that a coherent integration time or an incoherent integration time has reached a predetermined limit.
[0392] If it is determined at block 2915 that the calibration parameters have not reached one or more limits, the process continues directly to block 2917. However, if it is determined at block 2915 that the calibration parameters have reached one or more limits, the process continues to block 2916. In some alternative examples, block 2916 may involve scheduling a coordinated gap (e.g., for the next time block) in which content is not played by any of the coordinated audio devices and only one coordinated audio device plays the acoustic calibration signal. In some alternative examples, block 2916 may involve playing the content and the acoustic calibration signal by only one coordinated audio device. In other examples, block 2916 may involve playing the content by all coordinated audio devices and playing the acoustic calibration signal by only one coordinated audio device.
[0393] In this example, block 2917 involves assigning a calibration sequence for the next block in the current band. Block 2917 may, in some instances, involve increasing or decreasing the number of acoustic calibration signals that are simultaneously played during the next time block in the current frequency band. Block 2917 may involve, for example, determining when the last acoustic calibration signal was successfully demodulated in the current frequency band as part of the process of determining whether to increase or decrease the number of acoustic calibration signals that are simultaneously played during the next time block in the current frequency band.
[0394] FIG. 30 shows an example of a time-frequency allocation of calibration signals, gaps for noise estimation, and gaps for listening to a single audio device. FIG. 30 is intended to represent a snapshot in time of a continuous process, with various channel conditions present in each frequency band prior to time block 1. As with other disclosed examples, in FIG. 30, time is represented as a series of blocks represented along the horizontal axis, and frequency bands are represented along the vertical axis. The rectangles in FIG. 30 denoting "Device 1," "Device 2," etc. correspond to calibration signals for dominated audio device 1, dominated audio device 2, etc., in a particular frequency band during one or more time blocks.
[0395] The calibration signal in Band 1 (frequency band 1) essentially represents a repeated one-shot measurement for one time block. During each time block, except for time block 1, where a controlled gap is punched, there is a calibration signal for only one controlled audio device in Band 1.
[0396] In Band 2, during each time block, there are calibration signals for two coordinated audio devices. In this example, the calibration signals are assigned orthogonal codes. This configuration allows all coordinated audio devices to play their acoustic calibration signals in half the time required for the arrangement shown in Band 1. The calibration sequences for Devices 1 and 2 are completed by the end of Block 1, allowing a scheduled gap to be played in Block 2, which delays the playback of acoustic calibration signals by Devices 3 and 4 until Time Block 3.
[0397] In band 3, following potentially good conditions prior to time block 1, in the first block, four coordinated audio devices attempt to play their acoustic calibration signals. However, this causes poor demodulation results, so the concurrency is reduced to two devices in time block 2 (e.g., in block 2917 of FIG. 29). However, poor demodulation results are still returned. After the forced gap in time block 3, instead of further reducing the concurrency to a single device, starting in time block 4, longer codes are assigned to devices 1 and 2 in an attempt to improve robustness.
[0398] Band 4 begins with only Device 1 playing its acoustic calibration signal (e.g., via a four-block code sequence) during time blocks 1-4, possibly following poor conditions prior to time block 1. The code sequence is incomplete in block 4, where a gap is scheduled, delaying implementation of the forced gap by one time block.
[0399] The scenario depicted for band 5 proceeds much the same as the band 2 scenario, with two coordinated audio devices simultaneously playing their acoustic calibration signals during a single time block. In this example, the gap scheduled for time block 5 is delayed to time block 6 due to the delayed gap in band 4. Because, in this example, the minimum spectral spacing K G , two neighboring spectral blocks are not allowed to have simultaneous forced gaps.
[0400] FIG. 31 illustrates an audio environment, which in this example is a living space. As with other figures provided herein, the types, numbers, and arrangements of elements shown in FIG. 31 are provided merely as examples. Other implementations may include more, fewer, and / or different types, numbers, and / or arrangements of elements. In other examples, the audio environment may be another type of environment, such as an office environment, a vehicle environment, a park or other outdoor environment, etc. In this example, the elements in FIG. 31 include: 3101: A person, sometimes called a "user" or "listener"; 3102: A smart speaker including one or more loudspeakers and one or more microphones; 3103: A smart speaker including one or more loudspeakers and one or more microphones; 3104: A smart speaker including one or more loudspeakers and one or more microphones; 3105: A smart speaker including one or more loudspeakers and one or more microphones; 3106: Sound source 3106, which may be a noise source, located in the same room of the audio environment in which person 3101 and smart speakers 3102-3106 are located and has a known location. In some examples, sound source 3106 may be a legacy device, such as a radio, that is not part of the audio system including smart speakers 3102-3106. In some cases, the volume of sound source 3106 may not be continuously adjustable by person 3101 or adjustable by a command device. For example, the volume of sound source 3106 may be adjustable only by a manual process, such as via an on / off switch or by selecting a power or speed level (e.g., the power or speed level of a fan or air conditioner); 3107: Sound source 3107 may be a noise source that is not located in the same room of the audio environment as person 3101 and smart speakers 3102-3106. In some instances, sound source 3107 may not have a known location. In some instances, sound source 3107 may be diffuse.
[0401] The following description involves several basic assumptions. For example, it is assumed that estimates of the positions of the audio devices (such as smart devices 102-105 in FIG. 31) and estimates of the listener position (such as the position of person 101) are available. It is further assumed that a measure of interaudibility between the audio devices is known. This measure of interaudibility may, in some examples, be in the form of received levels in multiple frequency bands. Some examples are described below. In other examples, the measure of interaudibility may be a wideband measure, such as a measure that includes only one frequency band.
[0402] Readers may question whether microphones in consumer devices provide a uniform response, since mismatched microphone gains add a layer of ambiguity. However, the majority of smart speakers contain microelectromechanical system (MEMS) microphones, which are exceptionally well matched (within ±3 dB at worst, but typically within ±1 dB) and have a finite set of acoustic overload points. Therefore, an absolute mapping from digital dBFS (decibels relative to full scale) to dBSPL (decibels sound pressure level) can be determined by the model number and / or device descriptor. Therefore, MEMS microphones can be assumed to provide a well-calibrated acoustic standard for interaudibility measurements.
[0403] Figures 32, 33, and 34 are block diagrams representing three types of disclosed implementations. Figure 32 represents an implementation involving estimating the audibility (in this example, audibility in dBSPL) of all audio devices in an audio environment (e.g., the locations of smart speakers 3102-3105) at a user location (e.g., the location of person 3101 in Figure 31) based on the interaudibility between the audio devices, their physical locations, and the user's location. Such implementations do not require the use of a reference microphone at the user location. In some such examples, the audibility may be normalized by the digital level (in this example, in dBFS) of the loudspeaker drive signal to yield a transfer function between each audio device and the user. According to some examples, the implementation represented by Figure 32 is essentially a sparse interpolation problem: applying a model to estimate the level received at the listener location given banded levels measured between a set of audio devices at known locations.
[0404] In the example shown in FIG. 32, a full-matrix spatial audibility interpolator is shown receiving device geometry information (audio device position information), an interaudibility matrix (an example of which is provided below), and user position information, and outputting an interpolated transfer function. In this example, the interpolated transfer function is dBFS to dBSPL, which may be useful for leveling and equalizing audio devices such as smart devices. In some examples, there may be some null rows or columns in the audibility matrix corresponding to input-only or output-only devices. Implementation details corresponding to the example of FIG. 32 are provided below in "Full-Matrix Interaudibility Implementation."
[0405] FIG. 33 depicts an implementation involving estimating the audibility (in this example, in dBSPL) of an uncontrolled point source (such as sound source 3106 in FIG. 31) at a user location based on the audibility of the uncontrolled point source at the audio device, the physical location of the audio device, the location of the uncontrolled point source, and the location of the user. In some examples, the uncontrolled point source may be a noise source located in the same room as the audio device and person. In the example shown in FIG. 33, a point source spatial audibility interpolator is shown receiving device geometry information (audio device location information), an audibility matrix (an example of which is provided below), and sound source location information, and outputting interpolated audibility information.
[0406] FIG. 34 illustrates an implementation involving estimating the audibility (in this example, in dBSPL units) at a user location of a diffuse and / or unlocated, uncontrolled source (such as sound source 3107 in FIG. 31 ) based on the audibility of the sound source at each of the audio devices, the physical location of the audio devices, and the user's location. In this implementation, the location of the sound source is assumed to be unknown. The example shown in FIG. 34 illustrates a naive spatial audibility interpolator that receives device geometry information (audio device location information) and an audibility matrix (an example of which is described below) and outputs interpolated audibility information. In some examples, the interpolated audibility information referenced in FIGS. 3B and 3C may indicate interpolated audibility in dBSPL units, which may be useful for estimating the received level from a sound source (e.g., from a noise source). By interpolating the received level of a noise source, noise compensation (e.g., the process of increasing the gain of content in bands where noise is present) may be applied more accurately than can be achieved by referring to noise detected by a single microphone.
[0407] Full matrix interaudience implementation Table 5 shows what the terms in the equations in the following discussion represent. [Table 5]
[0408] L is the total number of audio devices, each of which has M i Let K be the total number of spectral bands reported by these audio devices. According to this example, we have a mutual audibility matrix H∈R that contains the measured transfer functions between all devices in all bands in linear units. K×L×L is determined.
[0409] There are several examples for determining H. However, the disclosed implementation is agnostic to the method used to determine H.
[0410] Some examples of determining H may involve multiple sequential iterations of "one-shot" calibrations using controlled acoustic calibration signals such as a swept sine wave, noise (e.g., white or pink noise), an acoustic DSSS signal, or curated program material played in turn by each of the audio devices. In some such examples, determining H may involve a sequential process of having a single smart audio device emit sound while other smart audio devices "listen" for sound.
[0411] 31, one such process may involve: (a) causing audio device 3102 to emit a sound and receiving microphone data corresponding to the emitted sound from the microphone arrays of audio devices 3103-3105; then (b) causing audio device 3103 to emit a sound and receiving microphone data corresponding to the emitted sound from the microphone arrays of audio devices 3102, 3104, and 3105; then (c) causing audio device 3104 to emit a sound and receiving microphone data corresponding to the emitted sound from the microphone arrays of audio devices 3102, 3103, and 3105; and then (d) causing audio device 3105 to emit a sound and receiving microphone data corresponding to the emitted sound from the microphone arrays of audio devices 3102, 3103, and 3104. These emitted sounds may or may not be the same, depending on the particular implementation.
[0412] Some pervasive and / or continuous methods involving acoustic calibration signals described in detail herein involve simultaneous playback of the acoustic calibration signal by multiple audio devices in an audio environment. In some such examples, the acoustic calibration signal is mixed into the played audio content. According to some implementations, the acoustic calibration signal is subaudible. Some such examples also include spectral hole punching (also referred to herein as "gap" formation).
[0413] According to some implementations, audio devices that include multiple microphones may estimate multiple audibility matrices (e.g., one per microphone), which are averaged to provide a single audibility matrix for each device. In some examples, anomalous data that may be due to a malfunctioning microphone may be detected and removed.
[0414] As mentioned above, the spatial position x of the audio device in 2D or 3D coordinates i It is assumed that the spatial location of the audio device, x, is also available. Several examples for determining device location based on time of arrival (TOA), direction of arrival (DOA), and a combination of DOA and TOA are described below. In other examples, the spatial location of the audio device, x, is determined based on the time of arrival (TOA), direction of arrival (DOA), and a combination of DOA and TOA. i may be determined by manual measurement, for example with a tape measure.
[0415] Additionally, the user's position x u is also assumed to be known, and in some cases both the position and orientation of the user may be known. Several methods for determining the listener position and listener orientation are described in detail below. According to some examples, the device position X=[x1x2...x L ] T is x u may be translated so that is at the origin of the coordinate system.
[0416] According to some implementations, the objective is to estimate an interpolated interaudibility matrix B by applying an appropriate interpolation to the measured data. In one example, a decay law model of the following form may be chosen:
number
[0417] In this example, x i represents the location of the transmitting device, and x j represents the location of the receiving device, and g i (k) represents the unknown linear output gain in band k, and α i (k) represents the distance attenuation constant. Least squares solution
number
number
number
[0418] In some embodiments
number
number
[0419] Figure 35 shows an example of a heatmap. In this example, heatmap 3500 represents the estimated transfer function for one frequency band from a sound source (o) to any point in a room with x and y dimensions as shown in Figure 35. The estimated transfer function is based on interpolation of measurements of the sound source by four receivers (x). The interpolated levels are calculated for any user position x in the room. u The heat map 3500 for
[0420] In another example, the distance decay model may include a critical distance parameter, such that the interpolation takes the form:
number
[0421] In this example, c i In some cases, the global room parameters d c represents a critical distance which may be solved as and / or constrained to be within a fixed range of values.
[0422] FIG. 36 is a block diagram illustrating an example of another implementation. As with other figures provided herein, the types, numbers, and arrangements of elements shown in FIG. 36 are provided merely as examples. Other implementations may include more, fewer, and / or different types, numbers, and / or arrangements of elements. In this example, the full matrix spatial audibility interpolator 3605, the delay compensation block 3610, the equalization and gain compensation block 3615, and the flexible renderer block 3620 are implemented by an instance of the control system 160 of the device 150 described above with reference to FIG. 1B. In some implementations, the device 150 may be a master device for an audio environment. According to some examples, the device 150 may be one of the audio devices of the audio environment. In some instances, the full matrix spatial audibility interpolator 3605, the delay compensation block 3610, the equalization and gain compensation block 3615, and the flexible renderer block 3620 may be implemented via instructions (e.g., software) stored on one or more non-transitory media.
[0423] In some examples, the full matrix spatial audibility interpolator 3605 may be configured to calculate the estimated audibility at the listener's position, as described above. According to this example, the equalization and gain compensation block 3615 calculates the interpolated audibility frequency band B i (k) Based on 3607, the equalization and compensation gain matrix 3617 (G∈R in Table 5) is K×LThe equalization and compensation gain matrix 3617 may be configured to determine a target curve (denoted as ). The equalization and compensation gain matrix 3617 may, in some cases, be determined using standardized techniques. For example, the estimated level at the user position may be smoothed across frequency bands, and equalization (EQ) gains may be calculated so that the result matches a target curve. In some implementations, the target curve may be spectrally flat. In other examples, the target curve may roll off gradually toward high frequencies to avoid overcompensation. In some cases, the EQ frequency bands may then be mapped to a different set of frequency bands that corresponds to the capabilities of a particular parametric equalizer. In some examples, the different set of frequency bands may be the 77 CQMF bands referred to elsewhere herein. In other examples, the different set of frequency bands may include a different number of frequency bands, for example, 20 critical bands, or as few as two frequency bands (high and low). Some implementations of the flexible renderer may use 20 critical bands.
[0424] In this example, the process of applying compensation gain and EQ is split so that compensation gain provides a coarse overall level match and EQ provides finer control in multiple bands. According to some alternative implementations, compensation gain and EQ may be implemented as a single process.
[0425] In this example, the flexible renderer block 3620 is configured to render audio data of the program content 3630 according to corresponding spatial information (e.g., position metadata) of the program content 3630. The flexible renderer block 3620 may be configured to implement CMAP, FV, a combination of CMAP and FV, or another type of flexible rendering, depending on the particular implementation. According to this example, the flexible renderer block 3620 is configured to use an equalization and compensation gain matrix 3617 to ensure that each loudspeaker is heard by the user at the same level with the same equalization. The loudspeaker signals 3625 output by the flexible renderer block 3620 may be provided to audio devices of an audio system.
[0426] According to this implementation, the delay compensation block 3610 calculates the delay according to the audio device geometry information and the user positioning information (in some examples, τ∈R in Table 1). L×1 The flexible renderer block 3620 is configured to determine delay compensation information 3612 (which may be or include a delay compensation vector denoted as ≡ ...
[0427] FIG. 37 is a flow diagram outlining one example of another method that may be performed by an apparatus or system such as those disclosed herein. The blocks of method 3700, as with other methods described herein, are not necessarily performed in the order presented. Furthermore, such methods may include more or fewer blocks than those shown and / or described. The blocks of method 3700 may be performed by one or more devices, which may be (or may include) a control system, such as control system 160 shown in FIG. 1B and described above, or one of the other disclosed control system examples. According to some examples, the blocks of method 3700 may be implemented by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media.
[0428] In this implementation, block 3705 involves causing, by the control system, multiple audio devices in the audio environment to play audio data. In this example, each audio device of the multiple audio devices includes at least one loudspeaker and at least one microphone. However, in some such examples, the audio environment may include at least one output-only audio device that has at least one loudspeaker but no microphone. Alternatively or additionally, in some such examples, the audio environment may include one or more input-only audio devices that have at least one microphone but no loudspeaker. Some examples of method 3700 in such a context are described below.
[0429] According to this example, block 3710 involves determining, by the control system, audio device location data including an audio device location for each audio device of the plurality of audio devices. In some examples, block 3710 may involve determining the audio device location data by referencing previously obtained audio device location data stored in memory (e.g., memory system 165 of FIG. 1B). In some instances, block 3710 may involve determining the audio device location data via an audio device automatic location process. The audio device automatic location process may involve performing one or more audio device automatic location methods, such as the DOA-based and / or TOA-based audio device automatic location methods referenced elsewhere herein.
[0430] According to this implementation, block 3715 involves obtaining, by the control system, microphone data from each audio device of the plurality of audio devices, in this example, the microphone data corresponding at least in part to sounds reproduced by loudspeakers of other audio devices in the audio environment.
[0431] In some examples, having multiple audio devices play audio data may involve having each audio device of the multiple audio devices play audio when all other audio devices in the audio environment are not playing audio. 31, one such process may involve: (a) causing audio device 3102 to emit a sound and receiving microphone data corresponding to the emitted sound from the microphone arrays of audio devices 3103-3105; then (b) causing audio device 3103 to emit a sound and receiving microphone data corresponding to the emitted sound from the microphone arrays of audio devices 3102, 3104, and 3105; then (c) causing audio device 3104 to emit a sound and receiving microphone data corresponding to the emitted sound from the microphone arrays of audio devices 3102, 3103, and 3105; and then (d) causing audio device 3105 to emit a sound and receiving microphone data corresponding to the emitted sound from the microphone arrays of audio devices 3102, 3103, and 3104. These emitted sounds may or may not be the same, depending on the particular implementation.
[0432] Other examples of block 3715 may involve obtaining microphone data while content is being played by each of the audio devices. Some such examples may involve spectral hole punching (also referred to herein as forming “gaps”). Thus, some such examples may involve causing each audio device of the plurality of audio devices, by the control system, to insert one or more frequency range gaps in audio data being played by one or more loudspeakers of each audio device.
[0433] In this example, block 3720 involves determining, by the control system, for each audio device of the plurality of audio devices, the interaudibility of each audio device to each other audio device of the plurality of audio devices. In some implementations, block 3720 may involve determining an interaudibility matrix, for example, as described above. In some examples, determining the interaudibility matrix may involve a process of mapping decibels relative to full scale to decibels of sound pressure level. In some implementations, the interaudibility matrix may include measured transfer functions between each audio device of the plurality of audio devices. In some examples, the interaudibility matrix may include a value for each frequency band of a plurality of frequency bands.
[0434] According to this implementation, block 3725 involves determining, by the control system, a user position of the person within the audio environment. In some examples, determining the user position may be based at least in part on at least one of direction of arrival data or time of arrival data corresponding to one or more utterances of the person. Some detailed examples of determining a user position of a person in an audio environment are described below.
[0435] In this example, block 3730 involves determining, by the control system, a user-position audibility of each audio device of the plurality of audio devices at the user position. According to this implementation, block 3735 involves controlling one or more aspects of audio device playback based at least in part on the user-position audibility. In some examples, the one or more aspects of audio device playback may include leveling and / or equalization, for example, as described above with reference to FIG. 36.
[0436] According to some examples, block 3720 (or another block of method 3700) may involve determining an interpolated interaudibility matrix by applying interpolation to the measured audibility data. In some examples, determining the interpolated interaudibility matrix may involve applying a decay law model based in part on a distance decay constant. In some examples, the distance decay constant may include per-device parameters and / or audio environment parameters. In some instances, the decay law model may be frequency band-based. According to some examples, the decay law model may include a critical distance parameter.
[0437] In some examples, method 3700 may involve estimating an output gain for each audio device of a plurality of audio devices according to values of an interaudibility matrix and an attenuation law model. In some instances, estimating the output gain for each audio device may involve determining a least-squares solution to a function of the values of the interaudibility matrix and the attenuation law model. In some examples, method 3700 may involve determining a value of an interpolated interaudibility matrix according to a function of the output gain for each audio device, the user position, and each audio device position. In some examples, the value of the interpolated interaudibility matrix may correspond to a user-position audibility for each audio device.
[0438] According to some examples, the method 3700 may involve equalizing frequency band values of an interpolated interaudibility matrix. In some examples, the method 3700 may involve applying a delay compensation vector to the interpolated interaudibility matrix.
[0439] As noted above, in some implementations, the audio environment may include at least one output-only audio device having at least one speaker but no microphone. In some such examples, method 3700 may involve determining the audibility of the at least one output-only audio device at the audio device location of each audio device of a plurality of audio devices.
[0440] As noted above, in some implementations, an audio environment may include one or more input-only audio devices that have at least one microphone but no loudspeaker. In some such examples, method 3700 may involve determining the audibility of each loudspeaker-equipped audio device in the audio environment at the respective locations of the one or more input-only audio devices.
[0441] Point noise source case implementation This section discloses an implementation corresponding to Figure 33. As used in this section, a "point noise source" is a point noise source at location x n A point noise source case refers to a noise source for which a source signal is available but no source signal is available. An example of this is when sound source 3106 in Figure 31 is the noise source. Instead of (or in addition to) determining an interaudibility matrix corresponding to the interaudibility of each of multiple audio devices in an audio environment, implementations of the "point noise source case" involve determining the audibility of such point source at each of multiple audio device locations. Some such examples use a noise audibility matrix A∈R that measures the received level of such point source at each of multiple audio device locations, rather than a transfer function as in the full matrix spatial audibility example described above. K×L Involved in determining
[0442] In some embodiments, the estimation of A may be performed in real time, for example, while the audio is playing in the audio environment. According to some implementations, the estimation of A may be part of a process to compensate for noise from point sources (or other sources of known location).
[0443] FIG. 38 is a block diagram illustrating an example of a system according to another implementation. As with other figures provided herein, the types, numbers, and arrangements of elements shown in FIG. 38 are provided merely as examples. Other implementations may include more, fewer, and / or different types, numbers, and / or arrangements of elements. According to this example, control systems 160A-160L correspond to audio devices 3801A-3801L (where L is 2 or greater) and are instances of control system 160 of apparatus 150 described above with reference to FIG. 1B. Here, control systems 160A-160L implement multi-channel acoustic echo cancellers 3805A-3805L.
[0444] In this example, the point source spatial audibility interpolator 3810 and the noise compensation block 3815 are implemented by the control system 160M of device 3820, which is another instance of device 150 described above with reference to FIG. 1B . In some examples, device 3820 may be what is referred to herein as an orchestration device or a smart home hub. However, in alternative examples, device 3820 may be an audio device. In some instances, the functionality of device 3820 may be implemented by one of audio devices 3801A-3801L. In some instances, the multi-channel acoustic echo cancellers 3805A-3805L, the point source spatial audibility interpolator 3810, and / or the noise compensation block 3815 may be implemented via instructions (e.g., software) stored on one or more non-transitory media.
[0445] In this example, sound source 3825 is generating sound 3830 in an audio environment. According to this example, sound 3830 is considered noise. In this case, sound source 3825 is not operating under the control of any of control systems 160A-160M. In this example, the location of sound source 3825 is known by control system 160M (in other words, provided and / or stored in memory accessible by control system 160M).
[0446] According to this example, multi-channel acoustic echo canceller 3805A receives microphone signal 3802A from one or more microphones of audio device 3801A and local echo reference 3803A corresponding to the audio being played by audio device 3801A. Multi-channel acoustic echo canceller 3805A is then configured to generate residual microphone signal 3807A (sometimes referred to as an echo-canceled microphone signal) and provide residual microphone signal 3807A to device 3820. In this example, residual microphone signal 3807A is assumed to correspond primarily to sound 3830 received at the location of audio device 3801A.
[0447] Similarly, multi-channel acoustic echo canceller 3805L receives microphone signal 3802L from one or more microphones of audio device 3801L and local echo reference 3803L corresponding to the audio being played by audio device 3801L. Multi-channel acoustic echo canceller 3805L is configured to output residual microphone signal 3807L to unit 3820. In this example, residual microphone signal 3807L is assumed to primarily correspond to sound 3830 received at the location of audio device 3801L. In some examples, multi-channel acoustic echo cancellers 3805A-3805L may be configured for echo cancellation in each of K frequency bands.
[0448] In this example, point source spatial audibility interpolator 3810 receives residual microphone signals 3807A-3807L, as well as audio device geometry (position data for each of audio devices 3801A-3801L) and source position data. According to this example, point source spatial audibility interpolator 3810 is configured to determine noise audibility information indicative of the received level of sound 3830 at each of the positions of audio devices 3801A-3801L. In some examples, the noise audibility information may include noise audibility data for each of K frequency bands, and in some instances may be calculated using the noise audibility matrix A∈R referenced above. K×L may be.
[0449] In some implementations, the point source spatial audibility interpolator 3810 (or another block of the control system 160M) may be configured to estimate noise audibility information 3812 indicative of the level of sound 3830 at a user position in the audio environment based on the user position data and the received level of sound 3830 at each of the positions of the audio devices 3801A-3801L. In some instances, estimating the noise audibility information 3812 may involve, for example, applying a distance attenuation model to estimate a noise level vector b∈R at the user position. K×1 The interpolation process may involve an interpolation process such as that described above by doing:
[0450] According to this example, noise compensation block 3815 is configured to determine a noise compensation gain 3817 based on the estimated noise level 3812 at the user location. In this example, noise compensation gain 3817 is a multi-band noise compensation gain (e.g., the noise compensation gain q∈R referenced above) that may vary depending on the frequency band. K×1For example, noise compensation gain may be higher in a frequency band corresponding to a higher estimated level of sound 3830 at the user's location. In some examples, noise compensation gain 3817 may be provided to audio devices 3801A-3801L, such that audio devices 3801A-3801L control playback of audio data according to noise compensation gain 3817. As suggested by dashed lines 3817A and 3817L, in some instances, noise compensation block 3815 may be configured to determine a noise compensation gain specific to each of audio devices 3801A-3801L.
[0451] FIG. 39 is a flow diagram outlining one example of another method that may be performed by an apparatus or system such as those disclosed herein. The blocks of method 3900, as with other methods described herein, are not necessarily performed in the order shown. Furthermore, such methods may include more or fewer blocks than those shown and / or described. The blocks of method 3900 may be performed by one or more devices, which may be (or may include) a control system such as that shown in FIG. 1B and described above, or one of the other disclosed control system examples. According to some examples, the blocks of method 3900 may be implemented by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media.
[0452] In this implementation, block 3905 involves receiving, by the control system, a residual microphone signal from each of a plurality of microphones in the audio environment. In this example, the residual microphone signal corresponds to sound from a noise source received at each of a plurality of audio device locations. In the example described above with reference to FIG. 38 , block 3905 involves control system 160M receiving residual microphone signals 3807A-3807L from multi-channel acoustic echo cancellers 3805A-3805L. However, in some alternative implementations, one or more of blocks 3905-3925 (and possibly all of blocks 3905-3925) may be performed by another control system, such as one of the audio device control systems.
[0453] According to this example, block 3910 involves obtaining, by the control system, audio device location data corresponding to each of a plurality of audio device locations, noise source location data corresponding to a location of a noise source, and user location data corresponding to a location of a person within the audio environment. In some examples, block 3910 may involve determining the audio device location data, noise source location data, and / or user location data by referencing previously obtained audio device location data stored in memory (e.g., memory system 115 of FIG. 1). In some instances, block 3910 may involve determining the audio device location data, noise source location data, and / or user location data via an automatic location process. The automatic location process may involve performing one or more automatic location methods, such as the automatic location methods referenced elsewhere herein.
[0454] According to this implementation, block 3915 involves estimating a noise level of a sound from a noise source at a user position based on the residual microphone signal, audio device position data, noise source position data, and user position data. In the example described above with reference to FIG. 38 , block 3915 may involve point source spatial audibility interpolator 3810 (or another block of control system 160M) estimating a noise level 3812 of sound 3830 at a user position within the audio environment based on the user position data and the received level of sound 3830 at each of the positions of audio devices 3801A-3801L. In some instances, block 3915 may involve, for example, applying a distance attenuation model to estimate a noise level vector b∈R at the user position. K×1 By estimating , the interpolation process may be engaged as described above.
[0455] In this example, block 3920 involves determining a noise compensation gain for each of the audio devices based on an estimated noise level of sound from a noise source at the user's location. In the example described above with reference to FIG. 38, block 3920 may involve noise compensation block 3815 determining noise compensation gain 3817 based on estimated noise level 3812 at the user's location. In some examples, the noise compensation gain may be a multi-band noise compensation gain (e.g., the noise compensation gain q∈R referenced above) that may vary depending on frequency band. K×1 ) may also be used.
[0456] According to this implementation, block 3925 involves providing a noise compensation gain to each of the audio devices. In the example described above with reference to Figure 38, block 3925 may involve apparatus 3820 providing noise compensation gains 3817A-3817L to each of audio devices 3801A-3801L.
[0457] Diffuse or unlocalized noise source implementation Locating a sound source, such as a noise source, is not always possible, especially when the sound sources are not located in the same room or when the sound source is highly occluded relative to the microphone array(s) that detects the sound. In such cases, estimating the noise level at the user's location can be viewed as a sparse interpolation problem with several known noise level values (e.g., one for each microphone or microphone array for each of multiple audio devices in the audio environment).
[0458] Such interpolation is performed using a general function f:R 2 →R. This can be expressed as 2 It represents the interpolation of known points in the kth band (represented by the term A) to an interpolated scalar value (represented by R). An example involves selecting subsets of three nodes (corresponding to the microphones or microphone arrays of three audio devices in an audio environment) to form a triangle of nodes, and solving for audibility within the triangle by bivariate linear interpolation. For any given node i, the received level in the kth band is calculated as A i (k) =ax i +by i +c. Solving for the unknowns gives
number
[0459] The interpolated audibility at any point within the triangle is:
number
[0460] Other examples may involve barycentric interpolation or cubic triangular interpolation, as described, for example, in "Noise Compensation Methods for Sound Sources in a Sound Field," in "Sound and Sound Compensation Methods for Sound Sources in a Sound Field," IEEE Transactions on Sound and Sound, Vol. 1, No. 1, pp. 111-114, which is incorporated herein by reference. Such interpolation methods are applicable to the noise compensation methods described above with reference to Figures 38 and 39, for example, by replacing the point-source spatial audibility interpolator 3810 of Figure 38 with a naive spatial interpolator implemented according to any of the interpolation methods described in this section, and by omitting the process of obtaining noise source position data in block 3910 of Figure 39. The interpolation methods described in this section do not provide spherical distance attenuation, but do provide a reasonable level of interpolation within the listening area. [Non-Patent Document 1] Amidror, Isaac, “Scattered data interpolation methods for electronic imaging systems: a survey,” Journal of Electronic Imaging Vol. 11, No. 2, April 2002, pp. 157-176.
[0461] Figure 40 shows an example floor plan for another audio environment, in this case a living space. As with other figures provided herein, the types and number of elements shown in Figure 40 are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements.
[0462] According to this example, environment 4000 includes a living room 4010 in the upper left, a kitchen 4015 in the lower center, and a bedroom 4022 in the lower right. Squares and circles distributed throughout the living space represent a set of loudspeakers 4005a-4005h, at least some of which may be smart speakers that, in some implementations, are conveniently positioned in the space but do not conform to a standardized layout (arbitrarily positioned). In some examples, a television 4030 may be configured, at least in part, to implement one or more disclosed embodiments. In this example, environment 4000 includes cameras 4011a-4011e distributed throughout the environment. In some implementations, one or more smart audio devices in environment 4000 may also include one or more cameras. The one or more smart audio devices may be single-purpose audio devices or virtual assistants. In some such examples, one or more cameras of optional sensor system 180 (see FIG. 1B ) may be present in or on television 4030, in a mobile phone, or in a smart speaker, such as one or more of loudspeakers 4005b, 4005d, 4005e, or 4005h. Cameras 4011a-4011e are not shown in all illustrations of environment 4000 presented in this disclosure, but each of environment 4000 may nevertheless include one or more cameras in some implementations.
[0463] Auto-localization of audio devices The assignee has produced several speaker localization techniques for cinemas and homes that are excellent solutions for the use cases for which they were designed. Some such methods are based on time-of-flight derived from the impulse response between the sound source and a microphone approximately co-located with each loudspeaker. System latency in the record and playback chain can also be estimated, but requires sample synchrony between clocks and the need for known test stimuli to estimate the impulse response.
[0464] Recent examples of sound source localization in this context relax the constraints by requiring intra-device microphone synchronization but not inter-device synchronization. Additionally, some such methods forgo the need to pass audio between sensors through low-bandwidth message passing, such as through detection of the time of arrival (TOA, also known as "time of flight") of direct (unreflected) sound or detection of the dominant direction of arrival (DOA) of direct sound. Each approach has several potential advantages and disadvantages. For example, some previously deployed TOA methods can determine device geometry excluding unknown translations, rotations, and reflections around one of three axes. With only one microphone per device, the rotation of individual devices is also unknown. Some previously deployed DOA methods can determine device geometry excluding unknown translations, rotations, and scaling. While some such methods can produce satisfactory results under ideal conditions, their robustness to measurement errors has not been demonstrated.
[0465] Some embodiments disclosed herein allow for localization of a collection of smart audio devices based on 1) the DOAs between each pair of audio devices in an audio environment and 2) minimization of a nonlinear optimization problem designed for input of data type 1). Other embodiments disclosed herein allow for localization of a collection of smart audio devices based on 1) the DOAs between each pair of audio devices in a system, 2) the TOAs between each pair of devices, and 3) minimization of a nonlinear optimization problem designed for input of data types 1) and 2).
[0466] FIG. 41 shows an example of the geometric relationships between four audio devices in an environment. In this example, audio environment 4100 is a room that includes television 4101 and audio devices 4105a, 4105b, 4105c, and 4105d. According to this example, audio devices 4105a-4105d are located at positions 1-4 in audio environment 4100, respectively. As with other examples disclosed herein, the types, numbers, locations, and orientations of elements shown in FIG. 41 are intended to be exemplary only. Other implementations may have different types, numbers, and arrangements of elements, such as more or fewer audio devices, audio devices in different locations, audio devices with different capabilities, etc.
[0467] In this implementation, each of audio devices 4105a-4105d is a smart speaker that includes a microphone system and a speaker system including at least one speaker. In some implementations, each microphone system includes an array of at least three microphones. According to some implementations, television 4101 may include a speaker system and / or a microphone system. In some such implementations, an auto-localization method may be used to automatically localize television 4101 or a portion of television 4101 (e.g., television speakers, television transceiver, etc.). This is described below with reference to audio devices 4105a-4105d, for example.
[0468] Some of the embodiments described in this disclosure allow for automatic localization of a set of audio devices, such as audio devices 4105a-4105d shown in FIG. 41, based on the direction of arrival (DOA) between each pair of audio devices, the time of arrival (TOA) of the audio signals between each pair of devices, or both the DOA and TOA of the audio signals between each pair of devices. In some cases, as in the example shown in FIG. 41, each of the audio devices is enabled with at least one driver unit and a microphone array, and the microphone array is capable of providing the direction of arrival of incoming sound. According to this example, double arrows 4110a-b represent sound transmitted by audio device 4105a and received by audio device 4105b, as well as sound transmitted by audio device 4105b and received by audio device 4105a. Similarly, double arrows 4110ac, 4110ad, 4110bc, 4110bd, and 4110cd represent sounds transmitted and received by audio device 4105a and audio device 4105c, sounds transmitted and received by audio device 4105a and audio device 4105d, sounds transmitted and received by audio device 4105b and audio device 4105c, sounds transmitted and received by audio device 4105b and audio device 4105d, and sounds transmitted and received by audio device 4105c and audio device 4105d, respectively.
[0469] In this example, each of the audio devices 4105a-4105d has an orientation represented by arrows 4115a-4115d, which may be defined in various ways. For example, the orientation of an audio device with a single loudspeaker may correspond to the direction in which the single loudspeaker is facing. In some examples, the orientation of an audio device with multiple loudspeakers facing in different directions may be indicated by the direction in which one of the loudspeakers is facing. In other examples, the orientation of an audio device with multiple loudspeakers facing in different directions may be indicated by the direction of a vector corresponding to the sum of the audio output in the different directions in which each of the multiple loudspeakers is facing. In the example shown in FIG. 41, the orientation of the arrows 4115a-4115d is defined with reference to a Cartesian coordinate system. In other examples, the orientation of the arrows 4115a-4115d may be defined with reference to another type of coordinate system, such as a spherical or cylindrical coordinate system.
[0470] In this example, the television 4101 includes an electromagnetic interface 103 configured to receive electromagnetic waves. In some examples, the electromagnetic interface 4103 may be configured to transmit and receive electromagnetic waves. According to some implementations, at least two of the audio devices 4105a-4105d may include an antenna system configured as a transceiver. The antenna system may be configured to transmit and receive electromagnetic waves. In some examples, the antenna system includes an antenna array having at least three antennas. Some of the embodiments described in this disclosure enable automatic localization of a set of devices, such as the audio devices 4105a-4105d and / or the television 101 shown in FIG. 1, based at least in part on the DOAs of the electromagnetic waves transmitted between the devices. Thus, double-headed arrows 4110ab, 4110ac, 4110ad, 4110bc, 4110bd, and 4110cd may also represent electromagnetic waves transmitted between the audio devices 4105a, 4105d.
[0471] According to some examples, the antenna system of a device (such as an audio device) may be co-located with a loudspeaker of the device, e.g., adjacent to the loudspeaker. In some such examples, the antenna system orientation may correspond to the loudspeaker orientation. Alternatively or additionally, the antenna system of the device may have a known or predetermined orientation relative to one or more loudspeakers of the device.
[0472] In this example, the audio devices 4105a-4105d are configured to wirelessly communicate with each other and other devices. In some examples, the audio devices 4105a-4105d may include a network interface configured for communication between the audio devices 4105a-4105d and other devices over the Internet. In some implementations, the auto-localization process disclosed herein may be performed by a control system of one of the audio devices 4105a-4105d. In other examples, the auto-localization process may be performed by another device in the audio environment 4100, such as a smart home hub, configured for wireless communication with the audio devices 4105a-4105d. In other examples, the auto-localization process may be performed at least in part by a device external to the audio environment 100, such as a server, based on information received from one or more of the audio devices 4105a-4105d and / or the smart home hub.
[0473] Figure 42 illustrates audio emitters located within the audio environment of Figure 41. Some implementations provide automatic localization of one or more audio emitters, such as person 4205 in Figure 42. In this example, person 4205 is at position 5. Here, sound emitted by person 4205 and received by audio device 4105a is represented by single arrow 4210a. Similarly, sound emitted by person 4205 and received by audio devices 4105b, 4105c, and 4105d are represented by single arrows 4210b, 4210c, and 4210d. The audio emitter may be localized based on the DOA of the audio emitter sound as captured by the audio devices 4105a-4105d and / or the television 4101, based on the difference in the TOA of the audio emitter sound as measured by the audio devices 4105a-4105d and / or the television 4101, or based on both the difference in DOA and TOA.
[0474] Alternatively or additionally, some implementations may provide for automatic loc...
Claims
1. causing a first audio device in the audio environment to generate a first calibration signal by the control system; causing the control system to insert the first calibration signal into a first audio playback signal corresponding to a first content stream to generate a first modified audio playback signal for the first audio device; causing the control system to cause the first audio device to play the first modified audio playback signal to generate a first audio device playback sound; causing a second audio device in the audio environment to generate a second calibration signal by the control system; causing the control system to insert the second calibration signal into a second audio playback signal corresponding to a second content stream to generate a second modified audio playback signal for the second audio device; causing the second audio device to reproduce the second modified audio playback signal by the control system to generate a second audio device playback sound; causing at least one microphone in the audio environment to detect, by the control system, at least the first audio device playback sound and the second audio device playback sound and generate a microphone signal corresponding to at least the first audio device playback sound and the second audio device playback sound; causing the control system to extract the first calibration signal and the second calibration signal from the microphone signal; causing the control system to estimate at least one first acoustic scene metric based at least in part on the first calibration signal and the second calibration signal; causing the control system to insert a first gap in a first frequency range of the first audio playback signal or the first modified audio playback signal during a first time interval of the first content stream, the first gap including an attenuation of the first audio playback signal in the first frequency range, and the first modified audio playback signal and the first audio device playback sound including the first gap; causing the control system to insert the first gap within the first frequency range of the second audio playback signal or the second modified audio playback signal during the first time interval, wherein the second modified audio playback signal and the second audio device playback sound include the first gap; causing the control system to extract audio data from the microphone signal in at least the first frequency range to generate extracted audio data of a calibration signal or unplayed sound without the presence of audio content; causing the control system to estimate at least one second acoustic scene metric based at least in part on the extracted audio data; controlling gap insertion and calibration signal generation based at least in part on a signal-to-noise ratio of a calibration signal of at least one audio device in at least one frequency band; Audio processing methods.
2. 2. The audio processing method of claim 1, wherein the first calibration signal corresponds to a first inaudible component of a sound reproduced by the first audio device, and the second calibration signal corresponds to a second inaudible component of a sound reproduced by the second audio device.
3. 2. The audio processing method of claim 1, wherein the first calibration signal comprises a first DSSS signal and the second calibration signal comprises a second DSSS signal.
4. 2. The audio processing method of claim 1, further comprising controlling gap insertion and calibration signal generation such that the calibration signal does not correspond to a gap time interval or a gap frequency range.
5. The audio processing method of claim 1 , further comprising controlling gap insertion and calibration signal generation based at least in part on the time since noise was estimated in at least one frequency band.
6. causing the target audio device to play the unmodified audio playback signal of the target device content stream to generate a target audio device playback sound; causing the control system to estimate at least one of a target audio device audibility or a target audio device location based at least in part on the extracted audio data; the unmodified audio playback signal includes audio content in the first gap; the microphone signal also corresponds to the target audio device playback sound; 2. The audio processing method of claim 1.
7. 7. The audio processing method of claim 6, wherein the unmodified audio playback signal does not include gaps inserted in any frequency range.
8. 2. The audio processing method of claim 1, wherein the at least one first acoustic scene metric further comprises one or more of time of flight, time of arrival, direction of arrival, range, audio device audibility, audio device impulse response, angle between audio devices, audio device position, and audio environment noise.
9. 2. The audio processing method of claim 1, wherein having the at least one first or second acoustic scene metric estimated comprises estimating the at least one first or second acoustic scene metric or having another device estimate the at least one first or second acoustic scene metric.
10. 10. The audio processing method of claim 1, further comprising controlling one or more aspects of audio device playback based at least in part on the at least one first or second acoustic scene metric.
11. 2. The audio processing method of claim 1, wherein a first content stream component of the sound reproduced by the first audio device causes perceptual masking of a first calibration signal component of the sound reproduced by the first audio device, and a second content stream component of the sound reproduced by the second audio device causes perceptual masking of a second calibration signal component of the sound reproduced by the second audio device.
12. 2. The audio processing method of claim 1, wherein the control system is a master device control system.
13. causing third through Nth audio devices in the audio environment to generate third through Nth calibration signals by the control system; causing the control system to insert the third through Nth calibration signals into the third through Nth content streams to generate third through Nth modified audio playback signals for the third through Nth audio devices; and causing the control system to cause the third through Nth audio devices to play corresponding instances of the third through Nth modified audio playback signals to generate third through Nth instances of audio device playback sounds.
2. The audio processing method of claim 1.
14. causing, by the control system, at least one microphone of each of the first through Nth audio devices to detect first through Nth instances of audio device reproduced sound and generate microphone signals corresponding to the first through Nth instances of audio device reproduced sound, wherein the first through Nth instances of audio device reproduced sound include the first audio device reproduced sound, the second audio device reproduced sound, and the third through Nth instances of audio device reproduced sound; causing the control system to extract the first through Nth calibration signals from the microphone signals, and wherein the at least one first acoustic scene metric is estimated based at least in part on the first through Nth calibration signals.
14. An audio processing method according to claim 13.
15. determining one or more calibration signal parameters for a plurality of audio devices in the audio environment, the one or more calibration signal parameters usable for generating a calibration signal; providing the one or more calibration signal parameters to each audio device of the plurality of audio devices.
2. The audio processing method of claim 1.
16. 16. The audio processing method of claim 15, wherein determining the one or more calibration signal parameters includes scheduling a time slot for each audio device of the plurality of audio devices for playing a modified audio playback signal, wherein a first time slot for a first audio device is different from a second time slot for a second audio device.
17. 16. The audio processing method of claim 15, wherein determining the one or more calibration signal parameters comprises determining a frequency band for each audio device of the plurality of audio devices to reproduce a modified audio reproduction signal.
18. 20. The audio processing method of claim 17, wherein the first frequency band for the first audio device is different from the second frequency band for the second audio device.
19. 16. The audio processing method of claim 15, wherein determining the one or more calibration signal parameters comprises determining a DSSS spreading code for each audio device of the plurality of audio devices.
20. 20. The audio processing method of claim 19, wherein the first spreading code for the first audio device is different from the second spreading code for the second audio device.
21. 20. The audio processing method of claim 19, further comprising determining at least one spreading code length based at least in part on the audibility of a corresponding audio device.
22. 16. The audio processing method of claim 15, wherein determining the one or more calibration signal parameters comprises applying an acoustic model based at least in part on the interaudibility of each of a plurality of audio devices in the audio environment.
23. determining that calibration signal parameters for an audio device are at a maximum robustness level; determining that a calibration signal from the audio device cannot be successfully extracted from the microphone signal; and causing all other audio devices to mute at least a portion of the corresponding audio device playback.
16. An audio processing method according to claim 15.
24. A system configured to perform the method of claim 1.
25. 10. One or more non-transitory media having stored thereon software including instructions for controlling one or more devices to perform the method of claim 1.
Citation Information
Patent Citations
Voice positioning using voice signal encoding and recognition
JP2014514791A
Dynamically adapting sound based on environmental characterization
US10511906B1
Audiolocation method and system combining use of audio fingerprinting and audio watermarking
US20150341890A1
Speaker device
WO2013190632A1
Forced gap insertion for pervasive listening
WO2020023856A1