Orchestration of acoustic direct-sequenced spread spectrum signals for acoustic scene metric estimation
DSSS signals are used to estimate acoustic scene metrics, addressing the challenge of optimizing audio device interaction and playback in multi-device environments by enhancing precision and audio quality.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- DOLBY LABORATORIES LICENSING CORP
- Filing Date
- 2021-12-02
- Publication Date
- 2026-05-18
AI Technical Summary
Existing audio systems lack efficient methods for estimating acoustic scene metrics, such as audibility and device positioning, which are crucial for optimizing audio playback and interaction in multi-device environments.
The implementation of direct sequence spread spectrum (DSSS) signals for audio devices to generate modified playback signals, which are detected and processed to estimate acoustic scene metrics like time of flight, range, and audibility, allowing for optimized audio device interaction and playback control.
Enables precise estimation of acoustic scene metrics, enhancing audio device coordination and playback performance in complex environments by accounting for device positioning and noise levels, thereby improving audio quality and interaction.
Smart Images

Figure 0007860985000040 
Figure 0007860985000041 
Figure 0007860985000042
Abstract
Description
[Technical Field]
[0001] Cross-reference with related applications This application claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 121,085 filed on 3 December 2020, U.S. Provisional Patent Application No. 63 / 260,953 filed on 7 September 2021, U.S. Provisional Patent Application No. 63 / 120,887 filed on 3 December 2020, and U.S. Provisional Patent Application No. 63 / 201,561 filed on 4 May 2021, the contents of which are incorporated herein by reference.
[0002] This disclosure relates to an audio processing system and method. [Background technology]
[0003] Audio devices and systems are widely used. Existing systems and methods for estimating acoustic scene metrics (e.g., the audibility of audio devices) are known, but improved systems and methods are desired.
[0004] Notation and Nomenclature Throughout this disclosure, including the claims, “speaker” and “loudspeaker” are used synonymously to refer to any acoustic emitting transducer (or set of transducers) driven by a single speaker feed. A typical headphone set includes two speakers. A speaker may be implemented to include multiple transducers (e.g., a woofer and a tweeter) driven by a single common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) may, in some cases, undergo different processing in different circuit branches connected to different transducers.
[0005] Throughout this disclosure, including the claims, the expression “perform” operations on a signal or data (e.g., filtering, scaling, transforming, or applying gain to a signal or data) is used broadly to mean performing operations directly on the signal or data, or performing operations on a processed version of the signal or data (e.g., a pre-filtered or pre-processed version of the signal before the operation is performed on it).
[0006] Throughout this disclosure, including the claims, the term “system” is used broadly to mean a device, system, or subsystem. For example, a subsystem that implements a decoder may be called a decoder system, and a system containing such a subsystem (for example, a system that generates X output signals in response to multiple inputs, where the subsystem generates M of the inputs and the other XM inputs are received from an external source) may also be called a decoder system.
[0007] Throughout this disclosure, including the claims, the term “processor” is used broadly to mean a system or device that is programmable or otherwise configurable (e.g., by software or firmware) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include field-programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipelined operations on audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.
[0008] Throughout this disclosure, including in the claims, the terms “couple” or “coupled” may be used to mean either a direct connection or an indirect connection. Therefore, when a first device is connected to a second device, the connection may be a direct connection or an indirect connection through other devices and connections.
[0009] In this specification, “smart device” generally refers to an electronic device configured to communicate with one or more other devices (or networks) via various wireless protocols such as Bluetooth, Zigbee, near-field communications, Wi-Fi, Light Fidelity (LiFi), 3G, 4G, and 5G, and capable of operating to some extent interactively and / or autonomously. Some representative types of smart devices include smartphones, smart cars, smart thermostats, smart doorbells, smart locks, smart refrigerators, phablets and tablets, smartwatches, smart bands, smart keychains, and smart audio devices. The term “smart device” may also refer to devices that exhibit some of the characteristics of ubiquitous computing, such as artificial intelligence.
[0010] In this specification, the term “smart audio device” is used to describe a smart device that is either a single-purpose audio device or a multi-purpose audio device (e.g., an audio device that implements at least some aspects of virtual assistant functionality). A single-purpose audio device is a device (e.g., a television) that includes or is connected to at least one microphone (and optionally includes or is connected to at least one speaker and / or at least one camera) and is generally or primarily designed to serve a single purpose. For example, a television can (or is thought to be able to) play audio from program material, but in most cases, a modern television runs some kind of operating system on which multiple applications, including an application for watching television, run locally. In this sense, a single-purpose audio device having one or more speakers and one or more microphones is often configured to run local applications and / or services for direct use of the speakers and microphones. There are also single-purpose audio devices that are configured to group audio across zones, or user-defined areas, to enable audio playback.
[0011] A single conventional multipurpose audio device implements at least some aspects of virtual assistant functionality, while other aspects of virtual assistant functionality may be implemented by one or more other devices, such as one or more servers, with which the multipurpose audio device is configured to communicate. Such a multipurpose audio device may be referred to herein as a “virtual assistant.” A virtual assistant is a device (e.g., a smart speaker or voice assistant integrated device) that includes or is connected to at least one microphone (and, if necessary, also includes or is connected to at least one speaker and / or at least one camera). In some examples, a virtual assistant may be cloud-enabled in a sense, or otherwise provide the ability to utilize multiple devices (different from the virtual assistant) for applications that are not fully implemented within or on the virtual assistant itself. In other words, at least some aspects of virtual assistant functionality, such as speech recognition functionality, may be implemented (at least partially) by one or more servers or other devices with which the virtual assistant can communicate over a network, such as the Internet. Multiple virtual assistants may collaborate, for example, in a highly discrete and conditionally defined manner. For example, two or more virtual assistants may collaborate in the sense that one of them (e.g., the one that is most certain to have heard the wake word) responds to that wake word. In some forms, multiple connected virtual assistants may form a kind of group managed by a single main application, which may be a virtual assistant (or may implement a virtual assistant).
[0012] In this specification, “wake word” is used broadly to mean any sound (e.g., a word uttered by a human, or any other sound). A smart audio device is configured to wake up in response to the detection of sound ("hearing") (using at least one microphone included in or connected to the smart audio device, or at least one other microphone). In this context, “awake” means that the device enters a state of waiting for a sound command (i.e., listening). In some examples, what may be called a “wake word” in this specification may include multiple words, such as a phrase.
[0013] In this specification, the term “wake word detector” refers to a device (or software containing instructions for configuring the device) configured to continuously explore the consistency between real-time sound (e.g., speech) features and a trained model. Typically, a wake word event is triggered whenever the wake word detector determines that the probability of a wake word being detected exceeds a predefined threshold. For example, the threshold may be a predetermined threshold adjusted to provide a good compromise between false acceptance and false rejection rates. After a wake word event, the device enters a state of listening for commands (sometimes called an “awakened” or “attentiveness” state), in which case it may pass received commands to larger, more computationally intensive recognizers.
[0014] In this specification, the terms “program stream” and “content stream” refer to a collection of one or more audio signals, and possibly video signals, of which at least a portion are intended to be heard together. Examples include music selections, movie soundtracks, movies, television programs, audio portions of television programs, podcasts, live voice calls, and synthesized voice responses from smart assistants. In some examples, a content stream may contain multiple versions of at least a portion of an audio signal, for example, the same dialogue in multiple languages. In such cases, only one version or a portion of the audio data (for example, a version corresponding to one language) is intended to be played at a time. [Overview of the Initiative] [Means for solving the problem]
[0015] At least some aspects of this disclosure may be implemented via one or more audio processing methods. In some examples, the methods(s) may be implemented at least partially by a control system and / or instructions (e.g., software) stored in one or more non-temporary media. Some methods include the control system causing a first audio device in an audio environment to generate a first direct sequence spread spectrum (DSSS) signal set. In some aspects, the control system may be or include an orchestration device control system. Some such methods include the control system causing the first DSSS signal set to insert into a first audio playback signal set corresponding to a first content stream to generate a first modified audio playback signal set for the first audio device. Some such methods include the control system causing the first audio device to play the first modified audio playback signal set to generate a first audio device playback sound.
[0016] Some such methods include causing the control system to cause a second audio device in the audio environment to generate a second DSSS signal set. Some such methods include causing the control system to insert the second DSSS signal set into a second content stream to generate a second modified audio playback signal set for the second audio device. Some such methods include causing the control system to play the second modified audio playback signal set for the second audio device to generate a second audio device playback sound. Some methods may include causing each of a plurality of audio devices in the audio environment to play the modified audio playback signal set simultaneously.
[0017] Some such methods include causing the control system to cause at least one microphone in the audio environment to detect at least the first audio device playback sound and the second audio device playback sound, and to generate a set of microphone signals corresponding to at least the first audio device playback sound and the second audio device playback sound. Some such methods include causing the control system to extract the first DSSS signal set and the second DSSS signal set from the microphone signal set. Some such methods include causing the control system to estimate at least one acoustic scene metric based at least partially on the first DSSS signal set and the second DSSS signal set. Some methods may include controlling one or more modes of audio device playback based at least partially on the at least one acoustic scene metric.
[0018] In some examples, the at least one acoustic scene metric may include one or more of time of flight, arrival time, range, audibility of an audio device, impulse response of an audio device, angle between audio devices, position of an audio device, noise of an audio environment, or signal-to-noise ratio. According to some examples, causing the at least one acoustic scene metric to be estimated may include estimating the at least one acoustic scene metric. Alternatively, or additionally, causing the at least one acoustic scene metric to be estimated may include causing at least one acoustic scene metric to be estimated by another device and may include the at least one acoustic scene metric.
[0019] In some examples, the first content stream component of the sound reproduced by the first audio device may cause perceptual masking of the first DSSS signal component of the sound reproduced by the first audio device. In some examples, the second content stream component of the sound reproduced by the second audio device may cause perceptual masking of the second DSSS signal component of the sound reproduced by the second audio device.
[0020] Some methods may include causing a control system to generate three or more direct sequence spectrum spread (DSSS) signals for three or more audio devices in the audio environment. Some such methods may include causing the control system to insert the three or more DSSS signals into three or more content streams and generate three or more groups of modified audio reproduction signals for the three or more audio devices. Some such methods may include causing the control system to reproduce corresponding instances of the three or more groups of modified audio reproduction signals on the three or more audio devices and generate three or more instances of the sound reproduced by the audio devices.
[0021] Some such methods may include causing the control system to generate third to Nth direct sequence spread spectrum (DSSS) signals for the third to Nth audio devices of the audio environment. Some such methods may include causing the control system to generate third to Nth modified audio playback signals for the third to Nth audio devices by inserting the third to Nth DSSS signals into third to Nth content streams. Some such methods may include causing the control system to cause the third to Nth audio devices to play corresponding instances of the third to Nth modified audio playback signals, thereby generating third to Nth instances of audio device playback sound.
[0022] Some methods may include causing the control system to cause at least one microphone of each of the first to Nth audio devices to detect first to Nth instances of audio device playback sound, thereby generating a group of microphone signals corresponding to the first to Nth instances of audio device playback sound. In some examples, the first to Nth instances of audio device playback sound may include the first audio device playback sound, the second audio device playback sound, and at least a third instance (in some examples, the third to Nth instances) of audio device playback sound.
[0023] Some such methods may include causing the control system to extract the first to Nth DSSS signals from the group of microphone signals. In some examples, the at least one acoustic scene metric may be estimated based at least in part on the first to Nth DSSS signals.
[0024] Some methods may involve determining one or more DSSS parameters for multiple audio devices in the audio environment. In some examples, the one or more DSSS parameters may be available for generating a set of DSSS signals. Some such methods may involve providing the one or more DSSS parameters to each of the multiple audio devices.
[0025] In some examples, determining the one or more DSSS parameters may include scheduling a time slot for each of the multiple audio devices to reproduce the modified audio playback signal set. In some such examples, the first time slot for the first audio device may be different from the second time slot for the second audio device.
[0026] In some examples, determining the one or more DSSS parameters may include determining the frequency band for reproducing the modified audio playback signal set for each of the plurality of audio devices. In some such examples, the first frequency band for the first audio device may be different from the second frequency band for the second audio device.
[0027] In some examples, determining the one or more DSSS parameters may include determining a spreading code for each of the multiple audio devices. According to some such examples, a first spreading code for a first audio device may be different from a second spreading code for a second audio device.
[0028] Some methods may involve determining at least one diffusion code length, based at least partially on the audibility of the corresponding audio devices. In some examples, determining the one or more DSSS parameters may involve applying an acoustic model, based at least partially on the interaudibility of each of the multiple audio devices in the audio environment.
[0029] According to some examples, determining the one or more DSSS parameters may include determining the current playback objective. Some such methods may include applying an acoustic model based at least partially on the interaudibility of each of the multiple audio devices in the audio environment to determine the estimated performance of the DSSS signal set in the audio environment. Some such methods may include determining the perceptual effect of the DSSS signal set in the audio environment by applying a perceptual model based on human sound perception. Some such methods may include determining the one or more DSSS parameters based at least partially on one or more of the current playback objective, the estimated performance, and the perceptual effect.
[0030] In some examples, determining the one or more DSSS parameters may include detecting a DSSS parameter change trigger. Some such methods may include determining one or more new DSSS parameters corresponding to a DSSS parameter change trigger. Some such methods may include providing the one or more new DSSS parameters to one or more audio devices in the audio environment.
[0031] According to some examples, detecting the DSSS parameter change trigger may include detecting one or more of the following: a new audio device in the audio environment, a change in the position of an audio device, a change in the orientation of an audio device, a change in the settings of an audio device, a change in the position of a person in the audio environment, a change in the type of audio content being played in the audio environment, a change in background noise in the audio environment, a change in the configuration of the audio environment including but not limited to a change in the configuration of doors or windows in the audio environment, a clock skew between two or more audio devices in the audio environment, a clock bias between two or more audio devices in the audio environment, a change in the interacousability between two or more audio devices in the audio environment, or a change in the purpose of playback.
[0032] Some methods may include processing the received microphone signal set to generate a pre-processed microphone signal set. In some such examples, the DSSS signal set may be extracted from the pre-processed microphone signal set. Processing the received microphone signal set may include, for example, applying one or more beamforming, bandpass filtering, or echo cancellation.
[0033] In some examples, extracting at least the first DSSS signal group and the second DSSS signal group from the microphone signal group may include applying a matched filter to the microphone signal group or a pre-processed version of the microphone signal group to generate a delay waveform group. In some examples, the delay waveform group may include at least a first delay waveform based on the first DSSS signal group and a second delay waveform based on the second DSSS signal group. In some methods, this may include applying a low-pass filter to the delay waveform group. In some examples, applying the matched filter may be part of the demodulation process. In some examples, the output of the demodulation process may be a demodulated coherent baseband signal.
[0034] Some methods may include estimating a bulk delay and providing the bulk delay estimate to the demodulation process. Some methods may include performing baseband processing on the demodulated coherent baseband signal. In some examples, the baseband processing may output at least one estimated acoustic scene metric.
[0035] In some examples, the baseband processing may include generating a non-coherently integrated delayed waveform based on a set of demodulated coherent baseband signals received during a non-coherent integration period. In some examples, generating the non-coherently integrated delayed waveform may include squaring the set of demodulated coherent baseband signals received during the non-coherent integration period to generate a set of squared demodulated baseband signals. In some such examples, this may include integrating the set of squared demodulated baseband signals. In some examples, the baseband processing may include applying one or more of the following processes to the non-coherently integrated delayed waveform: a leading edge estimation process, a steered response power estimation process, or a signal-to-noise estimation process.
[0036] Some methods may involve estimating the bulk delay. Some such examples may involve providing the bulk delay estimate to the baseband processing.
[0037] Some methods may include estimating at least a first noise power level at the location of a first audio device and estimating a second noise power level at the location of a second audio device. In some examples, the estimation of the first noise power level may be based on the first delay waveform, and the estimation of the second noise power level may be based on the second delay waveform. Some such examples may include generating a distributed noise estimate for the audio environment based at least in part on the estimated first noise power level and the estimated second noise power level.
[0038] Some methods may involve performing an asynchronous bidirectional ranging process to offset an unknown clock bias between two asynchronous audio devices. In some examples, the asynchronous bidirectional ranging process may be based on a set of DSSS signals transmitted by each of the two asynchronous audio devices. Some such examples may involve performing the asynchronous bidirectional ranging process between each of a plurality of audio device pairs in the audio environment.
[0039] Some methods may involve performing a clock bias estimation process to determine the estimated clock bias between two asynchronous audio devices. In some examples, the clock bias estimation process may be based on a set of DSSS signals transmitted by each of the two asynchronous audio devices. Some such examples may involve compensating for the estimated clock bias.
[0040] Some methods may include performing the clock bias estimation process between each of the multiple audio devices in the audio environment to generate multiple estimated clock biases. Some such examples may include compensating each of the multiple estimated clock biases.
[0041] Some methods may involve performing a clock skew estimation process to determine the estimated clock skew between two asynchronous audio devices. In some examples, the clock skew estimation process may be based on a set of DSSS signals transmitted by each of the two asynchronous audio devices. Some such examples may involve compensating for the estimated clock skew. Some methods may involve performing the clock skew estimation process between each of a plurality of audio devices in the audio environment to generate a plurality of estimated clock skews. Some such examples may involve compensating for each of the plurality of estimated clock skews.
[0042] Some methods may include detecting a DSSS signal transmitted by an audio device. In some examples, the DSSS signal may correspond to a first spreading code. Some such examples may include providing a second spreading code to the audio device. In some examples, the first spreading code may be, or include, a first pseudorandom number sequence reserved for a newly activated audio device.
[0043] In some examples, at least a portion of the first audio playback signal group, at least a portion of the second audio playback signal group, or at least a portion of each of the first and second audio playback signal groups corresponds to silence.
[0044] At least some aspects of this disclosure may be implemented via a device. For example, one or more devices may be capable of at least partially implementing the methods disclosed herein. In some aspects, the device may be or include an audio processing system having an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or a combination thereof.
[0045] In some embodiments, the apparatus may also have a loudspeaker system including at least one loudspeaker. In some embodiments, the apparatus may also have a microphone system including at least one microphone.
[0046] In some embodiments, the control system may be configured to receive a first content stream, which may include a first set of audio signals. In some such embodiments, the control system may be configured to render the first set of audio signals to generate a first set of audio playback signals. In some such embodiments, the control system may be configured to generate a first set of direct sequence spread spectrum (DSSS) signals. In some such embodiments, the control system may be configured to insert the first set of DSSS signals into the first set of audio playback signals to generate a first modified set of audio playback signals. In some embodiments, inserting the first set of DSSS signals into the first set of audio playback signals may include mixing the first set of DSSS signals with the first set of audio playback signals. In some such embodiments, the control system may be configured to cause the loudspeaker system to play the first modified set of audio playback signals to generate a first audio device playback sound.
[0047] In some examples, the control system may have a DSSS signal generator configured to generate a DSSS signal group. In some examples, the control system may have a DSSS signal modulator configured to modulate the DSSS signal group generated by the DSSS signal generator to generate a first DSSS signal group. In some examples, the control system may have a DSSS signal injector configured to insert the first DSSS signal group into a first audio playback signal group to generate a first modified audio playback signal group.
[0048] In some examples, the control system may be configured to receive from the microphone system a group of microphone signals corresponding to at least the first audio device playback sound and the second audio device playback sound. In some examples, the second audio device playback sound may correspond to a second group of modified audio playback signals reproduced by the second audio device. In some examples, the second group of modified audio playback signals may include a second group of DSSS signals. In some examples, the control system may be configured to extract at least the second group of DSSS signals from the group of microphone signals.
[0049] In some embodiments, the control system may be configured to receive from the microphone system a group of microphone signals corresponding to at least the first audio device playback sound and the second to Nth audio device playback sound microphone signals. In some examples, the second to Nth audio device playback sounds may correspond to the second to Nth modified audio playback signals reproduced by the second to Nth audio devices. In some examples, the second to Nth modified audio playback signals may include the second to Nth DSSS signals. In some embodiments, the control system may be configured to extract at least the second to Nth DSSS signals from the group of microphone signals.
[0050] In some examples, the control system may be configured to estimate at least one acoustic scene metric based at least partially on the second to nth DSSS signals. In some examples, the at least one acoustic scene metric may include one or more of the following: time of flight, time of arrival, range, audibility of the audio device, impulse response of the audio device, angle between audio devices, position of the audio device, noise of the audio environment, or signal-to-noise ratio. In some embodiments, the control system may be configured to control one or more modes of audio device playback based at least partially on the at least one acoustic scene metric and / or at least one audio device characteristic.
[0051] In some examples, the control system may be configured to determine one or more DSSS parameters for each of the multiple audio devices in the audio environment. In some examples, the one or more DSSS parameters may be available for generating a set of DSSS signals. In some such embodiments, the control system may be configured to provide the one or more DSSS parameters to each of the multiple audio devices.
[0052] In some examples, determining the one or more DSSS parameters may include scheduling a time slot for each of the multiple audio devices to reproduce the modified audio playback signal set. In some such examples, the first time slot for the first audio device may be different from the second time slot for the second audio device.
[0053] In some examples, determining the one or more DSSS parameters may include determining the frequency band for reproducing the modified audio playback signal set for each of the plurality of audio devices. In some examples, the first frequency band for the first audio device may be different from the second frequency band for the second audio device.
[0054] In some embodiments, determining the one or more DSSS parameters may include determining a diffusion code for each of the multiple audio devices. In some examples, a first diffusion code for a first audio device may be different from a second diffusion code for a second audio device. In some examples, the control system may be configured to determine at least one diffusion code length based at least in part on the audibility of the corresponding audio devices. In some embodiments, determining the one or more DSSS parameters may include applying an acoustic model based at least in part on the interaudibility of each of the multiple audio devices in the audio environment.
[0055] In some manner, determining the one or more DSSS parameters may include determining the current playback objective. In some such examples, determining the one or more DSSS parameters may include determining the estimated performance of the DSSS signal set in the audio environment by applying an acoustic model that is at least partially based on the interaudibility of each of the multiple audio devices in the audio environment. In some such examples, determining the one or more DSSS parameters may include determining the perceptual effect of the DSSS signal set in the audio environment by applying a perceptual model that is based on human auditory perception. In some such examples, determining the one or more DSSS parameters may be at least partially based on one or more of the current playback objective, the estimated performance, or the perceptual effect. In some examples, determining the one or more DSSS parameters may be at least partially based on the current playback objective, the estimated performance, and the perceptual effect.
[0056] In some embodiments, determining the one or more DSSS parameters may include detecting a DSSS parameter change trigger. In some such embodiments, the control system may be configured to determine one or more new DSSS parameters corresponding to the DSSS parameter change trigger. In some such embodiments, the control system may be configured to provide the one or more new DSSS parameters to one or more audio devices in the audio environment.
[0057] In some manner, detecting the DSSS parameter change trigger may include detecting one or more of the following: a new audio device in the audio environment, a change in the position of an audio device, a change in the orientation of an audio device, a change in the settings of an audio device, a change in the position of a person in the audio environment, a change in the type of audio content that can be played in the audio environment, a change in background noise in the audio environment, a change in the configuration of the audio environment including but not limited to a change in the configuration of doors or windows in the audio environment, a clock skew between two or more audio devices in the audio environment, a clock bias between two or more audio devices in the audio environment, a change in the interacousability between two or more audio devices in the audio environment, or a change in the purpose of playback.
[0058] In some embodiments, the control system may be configured to process the received microphone signal set to generate a pre-processed microphone signal set. In some such examples, the control system may be configured to extract a DSSS signal set from the pre-processed microphone signal set. In some embodiments, processing the received microphone signal set may include applying one or more of beamforming, bandpass filtering, or echo cancellation.
[0059] In some examples, extracting at least the second to nth DSSS signals from the microphone signal group may include applying a matched filter to the microphone signal group or a pre-processed version of the microphone signal group to generate the second to nth delay waveforms. In some such examples, the second to nth delay waveforms may correspond to each of the second to nth DSSS signals. In some examples, the control system may be configured to apply a low-pass filter to each of the second to nth delay waveforms.
[0060] In some embodiments, the control system may be configured to implement a demodulator. In some such embodiments, applying the matching filter may be part of the demodulation process performed by the demodulator. In some such examples, the output of the demodulation process may be a demodulated coherent baseband signal.
[0061] In some examples, the control system may be configured to estimate a bulk delay and provide the bulk delay estimate to the demodulator. In some embodiments, the control system may be configured to implement a baseband processor configured for baseband processing of the demodulated coherent baseband signal. In some such embodiments, the baseband processor may be configured to output at least one estimated acoustic scene metric.
[0062] In some examples, the baseband processing may include generating a non-coherently integrated delayed waveform based on a set of demodulated coherent baseband signals received during a non-coherent integration period. In some examples, generating the non-coherently integrated delayed waveform may include squaring the set of demodulated coherent baseband signals received during the non-coherent integration period to generate a set of squared demodulated baseband signals, and integrating the set of squared demodulated baseband signals. In some examples, the baseband processing may include applying one or more of the following processes to the non-coherently integrated delayed waveform: a leading-edge estimation process, a staired response power estimation process, or a signal-to-noise estimation process. In some examples, the control system may be configured to estimate a bulk delay and provide the bulk delay estimate to the baseband processor.
[0063] In some embodiments, the control system may be configured to estimate second to nth noise power levels at the locations of second to nth audio devices based on the second to nth delay waveforms. In some such examples, the control system may be configured to generate a distributed noise estimate for the audio environment based at least partially on the second to nth noise power levels.
[0064] In some examples, the control system may be configured to perform an asynchronous bidirectional ranging process to offset an unknown clock bias between two asynchronous audio devices. In some examples, the asynchronous bidirectional ranging process may be based on a set of DSSS signals transmitted by each of the two asynchronous audio devices. In some examples, the control system may be further configured to perform the asynchronous bidirectional ranging process between each of a plurality of audio device pairs in the audio environment.
[0065] In some embodiments, the control system may be configured to perform a clock bias estimation process for determining an estimated clock bias between two asynchronous audio devices. In some examples, the clock bias estimation process may be based on a set of DSSS signals transmitted by each of the two asynchronous audio devices. In some embodiments, the control system may be configured to compensate for the estimated clock bias.
[0066] In some examples, the control system may be configured to perform the clock bias estimation process between each of the multiple audio devices in the audio environment and generate a plurality of estimated clock biases. In some embodiments, the control system may be configured to compensate each of the plurality of estimated clock biases.
[0067] In some embodiments, the control system may be configured to perform a clock skew estimation process to determine the estimated clock skew between two asynchronous audio devices. In some embodiments, the clock skew estimation process may be based on a set of DSSS signals transmitted by each of the two asynchronous audio devices. In some such examples, the control system may be configured to compensate for the estimated clock skew.
[0068] In some examples, the control system may be configured to perform the clock skew estimation process between each of the multiple audio devices in the audio environment and generate a plurality of estimated clock skews. In some such examples, the control system may be configured to compensate for each of the plurality of estimated clock skews.
[0069] In some embodiments, the control system may be configured to provide a DSSS signal transmitted by the audio device. In some such embodiments, the DSSS signal may correspond to a first spreading code. In some such embodiments, the first spreading code may be, or include, a first pseudorandom number sequence reserved for the newly activated audio device. In some embodiments, the control system may be configured to provide a second spreading code for subsequent transmission to the audio device.
[0070] In some examples, the control system may be configured to simultaneously reproduce a set of modified audio playback signals to each of the multiple audio devices in the audio environment.
[0071] Some additional aspects of this disclosure may be implemented by one or more methods. In some examples, the methods (one or more) may be at least partially implemented by a control system and / or instructions (e.g., software) stored in one or more non-temporary media. Some methods may include the control system receiving a first content stream. The first content stream may include a first set of audio signals. Some such methods include the control system rendering the first set of audio signals to generate a first set of audio playback signals. Some such methods include the control system generating a first set of direct sequence spread spectrum (DSSS) signals. Some such methods include the control system generating a first modified set of audio playback signals by inserting the first set of DSSS signals into the first set of audio playback signals. Some such methods include the control system generating a first audio device playback sound by having a loudspeaker system play the first modified set of audio playback signals.
[0072] Some methods may include the control system receiving a group of microphone signals from the microphone system that correspond to at least the first audio device playback sound and the second audio device playback sound. In some examples, the second audio device playback sound may correspond to a second group of modified audio playback signals reproduced by the second audio device. In some examples, the second group of modified audio playback signals may include a second group of DSSS signals. Some methods may include the control system extracting at least the second group of DSSS signals from the group of microphone signals.
[0073] Some methods may include the control system receiving a group of microphone signals from the microphone system corresponding to at least the first audio device playback sound and the second to Nth audio device playback sounds. In some examples, the second to Nth audio device playback sounds may correspond to the second to Nth modified audio playback signals played by the second to Nth audio devices. In some examples, the second to Nth modified audio playback signals may include the second to Nth DSSS signals. Some methods may include the control system extracting at least the second to Nth DSSS signals from the group of microphone signals.
[0074] Some methods may involve the control system estimating at least one acoustic scene metric based at least partially on the second to nth DSSS signals. In some examples, the at least one acoustic scene metric includes one or more of the following: time of flight, time of arrival, range, audibility of an audio device, impulse response of an audio device, angle between audio devices, position of an audio device, noise of the audio environment, or signal-to-noise ratio.
[0075] Some methods may include controlling one or more modes of audio device playback by the control system based at least one acoustic scene metric, at least one audio device characteristic, or at least partially on both the at least one acoustic scene metric and the at least one audio device characteristic.
[0076] In some cases, the first content stream component of the sound played by the first audio device may cause perceptual masking of the first DSSS signal component of the sound played by the first audio device.
[0077] Some methods may involve the control system determining one or more DSSS parameters for each of the multiple audio devices in the audio environment. In some examples, the one or more DSSS parameters may be available for generating a set of DSSS signals. Some methods may involve the control system providing the one or more DSSS parameters to each of the multiple audio devices.
[0078] In some examples, determining the one or more DSSS parameters may include scheduling a time slot for each of the multiple audio devices to reproduce the modified audio playback signal set. In some examples, the first time slot for the first audio device may be different from the second time slot for the second audio device. According to some examples, determining the one or more DSSS parameters may include determining a frequency band for each of the multiple audio devices to reproduce the modified audio playback signal set. In some examples, the first frequency band for the first audio device may be different from the second frequency band for the second audio device.
[0079] In some examples, determining the one or more DSSS parameters may include determining a diffusion code for each of the multiple audio devices. In some examples, the first diffusion code for the first audio device may be different from the second diffusion code for the second audio device. In some examples, determining at least one diffusion code length may be determined, at least in part, based on the corresponding audio device audibility. In some examples, determining the one or more DSSS parameters may include applying an acoustic model that is at least in part based on the interaudibility of each of the multiple audio devices in the audio environment.
[0080] In some examples, at least a portion of the first group of audio signals may correspond to silence.
[0081] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices in accordance with instructions (e.g., software) stored in one or more non-temporary media. Such non-temporary media may include memory devices such as those described herein. Memory devices include, but are not limited to, random-access memory (RAM) devices and read-only memory (ROM) devices. Thus, some innovative aspects of the subject matter described herein can be implemented via one or more non-temporary media having stored software.
[0082] Details of one or more implementations of the subject matter described herein are illustrated in the accompanying drawings and the following description. Other features, embodiments, and advantages will become apparent from the specification, drawings, and claims. Note that the relative dimensions in the following figures may not be drawn to exact scale. [Brief explanation of the drawing]
[0083] Similar reference numerals and notations in various drawings indicate the same elements.
[0084] [Figure 1A] Figure 1A shows an example of an audio environment. [Figure 1B] Figure 1B is a block diagram showing examples of components of an apparatus capable of implementing various aspects of the present disclosure. [Figure 2] Figure 2 is a block diagram showing examples of audio device elements in several disclosed modes. [Figure 3] Figure 3 is a block diagram showing an example of an audio device element in another disclosed mode. [Figure 4] Figure 4 is a block diagram showing an example of an audio device element according to another disclosed embodiment. [Figure 5]Figure 5 is a graph showing examples of the levels of the content stream component and the DSSS signal component of the audio playback sound from an audio device, over a certain frequency range. [Figure 6] Figure 6 is a graph showing an example of the power of two DSSS signals with different bandwidths but the same center frequency. [Figure 7] Figure 7 shows the elements of an example orchestration module. [Figure 8] Figure 8 shows another example of an audio environment. [Figure 9] Figure 9 shows an example of the main lobes of the acoustic DSSS signal set generated by audio devices 100B and 100C in Figure 8. [Figure 10] Figure 10 is a graph showing an example of a time-domain multiple access (TDMA) scheme. [Figure 11] Figure 11 is a graph showing an example of a frequency domain multiple access (FDMA) scheme. [Figure 12] Figure 12 is a graph showing other examples of orchestration methods. [Figure 13] Figure 13 is a graph illustrating other examples of orchestration methods. [Figure 14] Figure 14 shows the elements of an audio environment in another example. [Figure 15] Figure 15 is a flowchart illustrating another example of the orchestration method for the disclosed audio device. [Figure 16] Figure 16 shows another example of an audio environment. [Figure 17] Figure 17 is a block diagram showing examples of DSSS signal demodulator elements, baseband processor elements, and DSSS signal generator elements in several disclosed modes. [Figure 18] Figure 18 shows the elements of a DSSS signal demodulator in another example. [Figure 19]Figure 19 is a block diagram showing examples of baseband processor elements in several disclosed modes. [Figure 20] Figure 20 shows an example of a delay waveform. [Figure 21] Figure 21 shows an example of a block in a different form. [Figure 22] Figure 22 shows another example of a block in a different form. [Figure 23] Figure 23 is a block diagram showing examples of audio device elements in several disclosed modes. [Figure 24] Figure 24 shows blocks representing other examples of configurations. [Figure 25] Figure 25 shows another example of an audio environment. [Figure 26] Figure 26 is a timing diagram for one example. [Figure 27] Figure 27 is a timing diagram showing the relevant clock terms when estimating the time of flight between two asynchronous audio devices in one example. [Figure 28] Figure 28 is a graph illustrating an example of detecting relative clock skew between two audio devices using a single acoustic DSSS signal. [Figure 29] Figure 29 is a graph illustrating an example of how relative clock skew between two audio devices is detected by measuring a single acoustic DSSS signal multiple times. [Figure 30] Figure 30 is a graph showing an example of an acoustic DSSS spreading code reserved for device discovery. [Figure 31] Figure 31 shows another example of an audio environment. [Figure 32A] Figure 32A shows an example of a set of delayed waveforms generated by the audio device 100C in Figure 31, based on a set of acoustic DSSS signals received from audio devices 100A and 100B. [Figure 32B] Figure 32B shows an example of a group of delayed waveforms generated by audio device 100B in Figure 31 based on the acoustic DSSS signals received from audio devices 100A and 100C. [Figure 33] Figure 33 is a flowchart illustrating another example of the disclosure method. [Figure 34] Figure 34 is a flowchart illustrating another example of the disclosure method. [Figure 35] Figure 35 is a flowchart illustrating examples of how multiple audio devices can coordinate a measurement session in several ways. [Figure 36A] Figure 36A is a flowchart illustrating examples of how multiple audio devices can coordinate a measurement session in several ways. [Figure 36B] Figure 36B is a flowchart illustrating examples of how multiple audio devices can coordinate a measurement session in several ways. [Modes for carrying out the invention]
[0085] To achieve compelling spatial reproduction of media and entertainment content, the physical layout and relative capabilities of available speakers must be evaluated and considered. Similarly, to provide high-quality voice interaction (both with virtual assistants and remote speakers), users need to both hear and be able to hear conversations played through loudspeakers. As collaborating devices are added to the audio environment, the devices will fall within a more commonly useful audio range, which is expected to increase overall usability for users. Increasing the number of speakers enhances the spatiality of media presentations, thus increasing immersion.
[0086] Such opportunities and experiences can potentially be realized with sufficient coordination and cooperation between devices. Acoustic information about each audio device is a crucial element of such coordination and cooperation. This acoustic information includes the audibility of each loudspeaker from various positions within the audio environment, as well as the amount of noise within the audio environment.
[0087] Some traditional methods for mapping and calibrating a group of smart audio devices require a dedicated calibration procedure, where known stimuli are played from audio devices (often one audio device at a time) while one or more microphones are recording. While this process can be appealing to some users due to its creative sound design, it becomes a barrier to widespread adoption because it must be repeated every time a device is added, removed, or even rearranged. Imposing such a procedure on the user can interfere with the normal operation of the device and may frustrate some users. Even more rudimentary methods include manual user intervention via software applications ("apps") and / or guided processes in which the user directs the physical location of audio devices within the audio environment. Such methods create further barriers to user adoption and may provide relatively less information to the system than a dedicated calibration procedure.
[0088] Generally, calibration and mapping algorithms require basic acoustic information for each audio device in an audio environment. Numerous methods have been proposed for this purpose, utilizing various basic acoustic measurement results and acoustic characteristics. Examples of acoustic characteristics obtained from microphone signals (also referred to herein as "acoustic scene metrics") for use in such algorithms include the following:
[0089] Estimated physical distance between devices (acoustic ranging) Estimated angle between devices (direction of arrival (DoA)) ○ Estimated impulse response between devices (e.g., using a swept sinusoidal stimulus or other measured signal) Estimated background noise
[0090] However, existing calibration and mapping algorithms are generally not implemented to respond to changes in the acoustic scene of an audio environment, such as human movement within the audio environment or the repositioning of audio devices within the audio environment.
[0091] This disclosure describes a technique involving a direct-sequenced spread spectrum (DSSS) signal injected into content being rendered by an audio device. Such a method can enable an audio device to generate observations after receiving signals transmitted by other audio devices in an audio environment. In some embodiments, each participating audio device in an audio environment may be configured to generate a DSSS signal, generate a modified audio playback signal by injecting the DSSS signal into a rendered loudspeaker feed signal, and generate a first audio device playback sound by having the loudspeaker system play the modified audio playback signal. In some embodiments, each participating audio device in an audio environment may be configured to do the above while detecting audio device playback sounds from other orchestrated audio devices in the audio environment and processing the audio device playback sounds to extract DSSS signals.
[0092] DSSS signals have traditionally been deployed in the context of telecommunications. When used in a telecommunications context, DSSS signals are used to spread transmitted data over a wider frequency range before it is transmitted through the channel to the receiver. In contrast, most or all of the disclosing modes do not involve the use of DSSS signals to modify or transmit data. Instead, these disclosing modes involve transmitting DSSS signals between audio devices in an audio environment. What happens to the transmitted DSSS signal between the transmitter and receiver is itself transmitted information. This is one major difference between how DSSS signals are used in a telecommunications context and how they are used in the disclosing modes.
[0093] Furthermore, the modes of disclosure relate to the transmission and reception of acoustic DSSS signals, rather than electromagnetic DSSS signals. In many modes of disclosure, acoustic DSSS signals are inserted into a content stream for rendering and playback, so that these acoustic DSSS signals are included in the played audio. According to some such modes, acoustic DSSS signals are inaudible to humans, so people in the audio environment do not perceive the acoustic DSSS signals and only perceive the played audio content.
[0094] Another difference between the use of acoustic DSSS signals disclosed herein and the use of DSSS signals in the context of telecommunications relates to what is referred to herein as the “near / far problem.” In some cases, the acoustic DSSS signals disclosed herein may be transmitted and received by numerous audio devices in an audio environment. The acoustic DSSS signals may overlap in time and frequency. Some of the disclosed modes depend on how DSSS spreading codes are generated to isolate the acoustic DSSS signals. In some cases, because the audio devices are in close proximity to each other, the signal levels may impair the isolation of the acoustic DSSS signals, making it difficult to isolate the signals. This is one manifestation of the near / far problem, and some solutions are disclosed herein.
[0095] Several methods may include receiving a first content stream containing a first set of audio signals; rendering the first set of audio signals to generate a first set of audio playback signals; generating a first set of direct sequence spread spectrum (DSSS) signals; generating a first set of modified audio playback signals by inserting the first set of DSSS signals into the first set of audio playback signals; and generating a first audio device playback sound by having the first set of modified audio playback signals played back by a loudspeaker system. One or more methods may include receiving a set of microphone signals corresponding to at least a first audio device playback sound and second to Nth audio device playback sounds corresponding to second to Nth modified audio playback signals (including second to Nth DSSS signals) played back by second to Nth audio devices; extracting second to Nth DSSS signals from the microphone signals; and estimating at least one acoustic scene metric based at least partially on the second to Nth DSSS signals.
[0096] The acoustic scene metrics (one or more) may be, or include, the audibility of audio devices, the impulse response of audio devices, the angle between audio devices, the position of audio devices, and / or noise in the audio environment. Some methods disclosed may include controlling one or more aspects of audio device playback based at least in part on the acoustic scene metrics (one or more).
[0097] Some of the methods disclosed may include arranging a plurality of audio devices to perform a method that includes DSSS signals. Some such methods may include, by a control system, causing a first audio device in an audio environment to generate a first set of DSSS signals; causing the control system to insert the first set of DSSS signals into a first set of audio playback signals corresponding to a first content stream to generate a first modified set of audio playback signals for the first audio device; and causing the control system to play the first modified set of audio playback signals on the first audio device to generate a first audio device playback sound.
[0098] Some such methods may include: causing a control system to cause a second audio device in the audio environment to generate a second DSSS signal group; causing the control system to insert the second DSSS signal group into a second content stream to generate a second modified audio playback signal group for the second audio device; causing the control system to play the second modified audio playback signal group for the second audio device to generate a second audio device playback sound; and causing the control system to play the second modified audio playback signal group for the second audio device to generate a second audio device playback sound.
[0099] Some such configurations may involve the control system causing at least one microphone in the audio environment to detect at least the first audio device playback sound and the second audio device playback sound, and to generate a set of microphone signals corresponding to at least the first audio device playback sound and the second audio device playback sound. Some such configurations may involve the control system extracting at least the first DSSS signal set and the second DSSS signal set from the microphone signal set, and the control system estimating at least one acoustic scene metric based at least partially on the first DSSS signal set and the second DSSS signal set.
[0100] Figure 1A shows an example of an audio environment. As with other figures provided herein, the types and number of elements shown in Figure 1A are given for illustrative purposes only. Other forms may include more, fewer, and / or different types and numbers of elements.
[0101] In this example, the audio environment 130 is the living space of a home. In the example shown in Figure 1A, audio devices 100A, 100B, 100C, and 100D are placed within the audio environment 130. In this example, each of the audio devices 100A to 100D includes a corresponding loudspeaker system 110A, 110B, 110C, and 110D. In this example, the loudspeaker system 110B of audio device 100B includes at least a left loudspeaker 110B1 and a right loudspeaker 110B2. In this example, the audio devices 100A to 100D include loudspeakers of various sizes and capabilities. At the time shown in Figure 1A, the audio devices 100A to 100D are producing corresponding instances of the audio device playback sounds 120A, 120B1, 120B2, 120C, and 120D.
[0102] In this example, each of the audio devices 100A to 100D includes a corresponding microphone system 111A, 111B, 111C, or 111D. Each of the microphone systems 111A to 111D includes one or more microphones. In some examples, the audio environment 130 may include at least one audio device lacking a loudspeaker system, or at least one audio device lacking a microphone system.
[0103] In some cases, at least one acoustic event may occur in the audio environment 130. For example, one such acoustic event may, in some cases, be caused by a speaker issuing a voice command. In other cases, the acoustic event may be caused, at least partially, by a variable element of the audio environment 130, such as a door or window. For example, when a door is opened, sounds from outside the audio environment 130 may be perceived more clearly inside the audio environment 130. Furthermore, a change in the angle of the door may alter part of the echo path within the audio environment 130.
[0104] Figure 1B is a block diagram showing examples of components of an apparatus capable of carrying out various embodiments of the present disclosure. As with other figures provided herein, the types and number of elements shown in Figure 1B are given merely as examples. In other embodiments, there may be more, fewer, and / or different types and numbers of elements. According to some examples, apparatus 150 may be configured to perform at least some of the methods disclosed herein. In some embodiments, apparatus 150 may be, or include, one or more components of an audio system. For example, in some embodiments, apparatus 150 may be an audio device such as a smart audio device. In other embodiments, apparatus 150 may be a mobile device (such as a cell phone), a laptop computer, a tablet device, a television, or other type of device.
[0105] In the example shown in Figure 1A, audio devices 100A to 100D are instances of apparatus 150. In some examples, the audio environment 100 in Figure 1A may include an orchestration device, such as what may be referred to herein as a smart home hub. The smart home hub (or other orchestration device) may be an instance of apparatus 150. In some embodiments, one or more of the audio devices 100A to 100D may function as an orchestration device.
[0106] In some alternative forms, device 150 may be a server or include one. In some such examples, device 150 may be an encoder or include one. Thus, in some cases, device 150 may be a device configured for use in an audio environment such as a home audio environment, while in other cases, device 150 may be a device configured for use in a “cloud,” such as a server.
[0107] In this example, the device 150 includes an interface system 155 and a control system 160. The interface system 155 may, in some embodiments, include a wired or wireless interface configured to communicate with one or more other devices in an audio environment. In some embodiments, the audio environment may be a home audio environment. In other embodiments, the audio environment may be other types of environments, such as an office environment, a car environment, a train environment, a road or sidewalk environment, or a park environment. In some embodiments, the interface system 155 may be configured to exchange control information and related data with audio devices in the audio environment. In some embodiments, the control information and related data may relate to one or more software applications running on the device 150.
[0108] The interface system 155 may be configured in several ways to receive or provide a content stream. The content stream may include audio data. The audio data may include, but is not limited to, a set of audio signals. In some examples, the audio data may include channel data and / or spatial data such as spatial metadata. The metadata may be provided, for example, by what may be called a “coder” herein. In some examples, the content stream may include video data and audio data corresponding to the video data.
[0109] The interface system 155 may include one or more network interfaces and / or one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). In some embodiments, the interface system 155 may include one or more wireless interfaces, such as WiFi or Bluetooth. TM It may include configurations for communications.
[0110] The interface system 155 may, in some examples, include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 155 may include one or more interfaces between the control system 160 and a memory system, such as the optional memory system 165 shown in Figure 1B. However, the control system 160 may include a memory system in some cases. In some manner, the interface system 155 may be configured to receive input from one or more microphones in the environment.
[0111] In some manner, the control system 160 may be configured to perform, at least in part, the methods disclosed herein. The control system 160 may include, for example, a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.
[0112] In some embodiments, the control system 160 may reside on multiple devices. For example, in some embodiments, part of the control system 160 may reside on a device within one of the environments described herein, and another part of the control system 160 may reside on a device outside the environment, such as a server or a mobile device (e.g., a smartphone or tablet computer). In other embodiments, part of the control system 160 may reside on a device within one of the environments described herein, and another part of the control system 160 may reside on one or more other devices in the environment. For example, the functions of the control system may be distributed among multiple smart audio devices in the environment, or shared by an orchestration device (such as one which may be referred to herein as a smart home hub) and one or more other devices in the environment. In other embodiments, part of the control system 160 may reside on a device implementing a cloud-based service, such as a server, and another part of the control system 160 may reside on another device implementing a cloud-based service, such as another server or memory device. The interface system 155 may also reside on multiple devices in some embodiments.
[0113] Some or all of the methods described herein may be performed by one or more devices in accordance with instructions (e.g., software) stored in one or more non-temporary media. Such non-temporary media may include memory devices such as those described herein. Such memory devices include, but are not limited to, random-access memory (RAM) devices and read-only memory (ROM) devices. One or more non-temporary media may reside, for example, in the optional memory system 615 and / or control system 160 illustrated in Figure 1B. Thus, various innovative embodiments of the subject matter described herein can be implemented in one or more non-temporary media having stored software. The software may include, for example, instructions for controlling at least one device that performs some or all of the methods disclosed herein. The software may be executable by one or more components of a control system, such as the control system 160 in Figure 1B.
[0114] In some examples, the device 150 may include an optional microphone system 111, as shown in Figure 1B. The optional microphone system 111 may include one or more microphones. According to some examples, the optional microphone system 111 may include an array of microphones. In some cases, the array of microphones may be configured to perform receiving beamforming, for example, according to instructions from the control system 160. In some examples, the array of microphones may be configured to determine direction of arrival (DOA) and / or time of arrival (TOA) information, for example, according to instructions from the control system 160. Alternatively or additionally, the control system 160 may be configured to determine direction of arrival (DOA) and / or time of arrival (TOA) information, for example, according to a set of microphone signals received from the microphone system 111.
[0115] In some embodiments, one or more microphones may be part of or associated with another device, such as a speaker in a speaker system or a smart audio device. In some examples, the device 150 may not include a microphone system 111. However, in some such embodiments, the device 150 may nevertheless be configured to receive microphone data from one or more microphones in the audio environment via the interface system 160. In some such embodiments, the cloud-based embodiment of the device 150 may be configured to receive microphone data, or data corresponding to microphone data, from one or more microphones in the audio environment via the interface system 160.
[0116] In some embodiments, the device 150 may include an optional loudspeaker system 110, as shown in Figure 1B. The optional loudspeaker system 110 may include one or more loudspeakers, also referred to herein as “speakers” or more generally as “audio playback transducers.” In some examples (e.g., cloud-based embodiments), the device 150 may not include the loudspeaker system 110.
[0117] In some embodiments, the device 150 may include an optional sensor system 180, as shown in Figure 1B. The optional sensor system 180 may include one or more touch sensors, gesture sensors, motion detectors, etc. In some embodiments, the optional sensor system 180 may include one or more cameras. In some embodiments, the cameras may be free-standing cameras. In some examples, one or more cameras of the optional sensor system 180 may be located in a smart audio device. The smart audio device may be a single-purpose audio device or a virtual assistant. In some such examples, one or more cameras of the optional sensor system 180 may be located in a television, mobile phone, or smart speaker. In some embodiments, the device 150 may not include the sensor system 180. However, in some such embodiments, the device 150 may nevertheless be configured to receive sensor data from one or more sensors in the audio environment via the interface system 160.
[0118] In some embodiments, the device 150 may include an optional display system 185, as shown in Figure 1B. The optional display system 185 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some examples, the optional display system 185 may include one or more organic light-emitting diode (OLED) displays. In some examples, the optional display system 185 may include one or more displays for a smart audio device. In other examples, the optional display system 185 may include a television display, a laptop display, a mobile device display, or other types of displays. In some embodiments where the device 150 includes the display system 185, the sensor system 180 may include a touch sensor system and / or a gesture sensor system adjacent to one or more displays of the display system 185. According to some such embodiments, the control system 160 may be configured to control the display system 185 to present one or more graphical user interfaces (GUIs).
[0119] In some such examples, device 150 may be or include a smart audio device. In some such forms, device 150 may be or include a wake word detector. For example, device 150 may be or include a virtual assistant.
[0120] Figure 2 is a block diagram showing examples of audio device elements in several disclosed modes. As with other figures provided herein, the types and number of elements shown in Figure 2 are given merely as examples. Other modes may include more, fewer, and / or different types and numbers of elements. In this example, audio device 100A in Figure 2 is an instance of the apparatus 150 described above with reference to Figure 1B. In this example, audio device 100A is one of several audio devices in an audio environment and may in some cases be an example of audio device 100A shown in Figure 1A. In this mode, audio device 100A is one of several orchestrated audio devices in an audio environment. In this example, the audio environment includes at least two other orchestrated audio devices, audio device 100B and audio device 100C.
[0121] In this configuration, the audio device 100A includes the following elements:
[0122] 110A: An instance of the loudspeaker system 110 of Figure 1B, which includes one or more loudspeakers. 111A: An instance of the microphone system 111 in Figure 1B, which includes one or more microphones. 120A, B, C: Audio device playback sounds corresponding to rendered content played back by audio devices 100A to 100C within the same acoustic space. 201A: Audio playback signals output by rendering module 210A. 202A: Modified audio playback signals output by DSSS signal injector 211A. 203A: A group of DSSS signals output by the DSSS signal generator 212A. 204A: A group of DSSS signal replicas corresponding to a group of DSSS signals generated by other audio devices in the audio environment (in this example, at least audio devices 100B and 100C). In some examples, the DSSS signal replica group 204A may come from an external source such as an orchestration device (which may be other audio devices in the audio environment, other local devices such as a smart home hub, etc.) (e.g., WiFi or Bluetooth). TM It may be received (via wireless communication protocols such as...)
[0123] 205A: DSSS information belonging to and / or used by one or more audio devices in an audio environment. DSSS information 205A may include parameters used by the control system 160 of audio device 100A for purposes such as generating a DSSS signal set, modulating a DSSS signal set, or demodulating a DSSS signal set. DSSS information 205A may include one or more DSSS spreading code parameters and one or more DSSS carrier parameters. DSSS spreading code parameters may include, for example, DSSS spreading code length information, chipping rate information (or chip period information), etc. One chip period is the time required for one chip (bit) of the spreading code to be reproduced. The reciprocal of the chip period is the chipping rate. A group of bits in a DSSS spreading code is sometimes called a "chip" to indicate that it does not contain data (like a normal bit). In some examples, DSSS spreading code parameters may include a pseudorandom sequence. In some examples, DSSS information 205A may indicate which audio device is generating the acoustic DSSS signal set. In some cases, the DSSS information 205A may be received from an external source, such as an orchestration device (for example, via wireless communication).
[0124] 206A: Microphone (single or multiple) - a group of microphone signals received by 111A.
[0125] 208A: Demodulated coherent baseband signal group.
[0126] 210A: A rendering module configured to render audio signals from content streams such as music, movies, and TV programs, and generate audio playback signals.
[0127] 211A: A DSSS signal injector configured to insert a group of DSSS signals 230A modulated by a DSSS signal modulator 220A into a group of audio playback signals generated by a rendering module 210A to generate a modified audio playback signal group. The insertion process may be, for example, a mixing process in which the group of DSSS signals 230A modulated by the DSSS signal modulator 220A is mixed with a group of audio playback signals generated by a rendering module 210A to generate a modified audio playback signal group.
[0128] 212A: A DSSS signal generator configured to generate a DSSS signal group 203A and supply the DSSS signal group 203A to a DSSS signal modulator 220A and a DSSS signal demodulator 214A. In this example, the DSSS signal generator 212A includes a DSSS spreading code generator and a DSSS carrier generator. In this example, the DSSS signal generator 212A supplies a DSSS signal replica group 204A to the DSSS signal demodulator 214A.
[0129] 214A: A DSSS signal demodulator configured to demodulate a group of microphone signals 206A received by one or more microphones 111A. In this example, the DSSS signal demodulator 214A outputs a demodulated coherent baseband signal group 208A. Demodulation of the microphone signal group 206A can be performed using standard correlation techniques, for example, a matched filtering correlator bank in the form of integrate and dump. Several detailed examples are given below. To improve the performance of these demodulation techniques, in some embodiments, the microphone signal group 206A may be filtered before demodulation to remove unwanted content / phenomena. According to some embodiments, the demodulated coherent baseband signal group 208A may be filtered before being fed to the baseband processor 218A. The signal-to-noise ratio (SNR) generally improves as the integration time increases (as the length of the spreading code used increases).
[0130] 218A: A baseband processor configured for baseband processing of the demodulated coherent baseband signal group 208A. In some examples, the baseband processor 218A may be configured to implement techniques such as non-coherent averaging to improve the SNR by reducing the variance of the squared waveform to generate a delayed waveform. Several detailed examples are shown below. In this example, the baseband processor 218A is configured to output one or more estimated acoustic scene metrics 225A.
[0131] 220A: A DSSS signal modulator configured to modulate the DSSS signal group 203A generated by the DSSS signal generator to produce the DSSS signal group 230A.
[0132] 225A: Observations derived from one or more DSSSs, also referred to herein as acoustic scene metrics. The acoustic scene metrics (singular or plural) 225A include, or may include, data corresponding to time of flight, time of arrival, range, audibility of audio devices, impulse response of audio devices, angle between audio devices, position of audio devices, noise of the audio environment, and / or signal-to-noise ratio.
[0133] 233A is an acoustic scene metric processing module configured to receive and apply the acoustic scene metric group 225A. In this example, the acoustic scene metric processing module 233A is configured to generate information 235A (and / or command group) at least partially based on at least one acoustic scene metric 225A and / or at least one audio device characteristic. The audio device characteristic(s) may correspond to audio device 100A or to another audio device in the audio environment, depending on the particular mode. The audio device characteristic(s) may be stored in the memory of the control system 160, for example, or may be accessible to the control system 160.
[0134] 235A: Information for controlling one or more aspects of audio processing and / or audio device playback. Information 235A may include, for example, information (and / or command sets) for controlling a rendering process, an audio environment mapping process (such as an automatic audio device positioning process), an audio device calibration process, a noise suppression process and / or an echo attenuation process.
[0135] Examples of acoustic scene metrics As described above, in some forms, the baseband processor 218A (or another module of the control system 160) may be configured to determine one or more acoustic scene metrics 225A. The following are some examples of the acoustic scene metric group 225A.
[0136] Distance measurement The DSSS signal received by an audio device from another device contains information about the distance between the two devices in the form of the signal's time-of-flight (ToF). Therefore, according to some examples, a control system may be configured to extract delay information from the demodulated DSSS signal and convert that delay information into a pseudo-range measurement, for example, as follows:
[0137]
number
[0138] In the aforementioned equation, τ represents delay information (also referred to as ToF herein), ρ represents the pseudo-range measurement, and c represents the speed of sound. The term "pseudo-range" is used because the range itself is not directly measured; therefore, the range between devices is estimated according to timing estimates. In a distributed asynchronous system of audio devices, each audio device operates on its own clock, resulting in bias in the raw delay measurement set. With a sufficient set of delay measurement values, these biases can be eliminated, and sometimes even estimated. Detailed examples of delay information extraction, generation and use of pseudo-range measurement sets, and determination and resolution of clock bias are provided below.
[0139] DoA Similar to ranging, using multiple microphones available on the listening device, the control system may be configured to estimate the direction of arrival (DoA) by processing a demodulated set of acoustic DSSS signals. In some such configurations, the resulting DoA information can be used as input to a DoA-based audio device automatic positioning method.
[0140] audible The signal intensity of the demodulated acoustic DSSS signal is proportional to the audibility of the audio device as heard in the frequency band from which the audio device transmits the acoustic DSSS signal set. In some embodiments, the control system can be configured to obtain a banded estimate for the entire frequency range by making multiple observations over a certain range of frequency bands. If the digital signal level of the transmitting audio device is known, the control system can, in some examples, be configured to estimate the absolute acoustic gain of the transmitting audio device.
[0141] Figure 3 is a block diagram showing an example of audio device elements in another disclosed mode. As with other figures provided herein, the types and number of elements shown in Figure 3 are given merely as examples. Other modes may include more, fewer, and / or different types and numbers of elements. In this example, audio device 100A in Figure 3 is an instance of the device 150 described above with reference to Figures 1B and 2. However, in this mode, audio device 100A is configured to organize multiple audio devices in an audio environment, including at least audio devices 100B, 100C, and 100D.
[0142] The configuration shown in Figure 3 includes several additional elements in addition to all the elements in Figure 2. Elements common to both Figure 2 and Figure 3 will not be explained again here, except that their function may differ in the configuration shown in Figure 3. According to this configuration, audio device 100A includes the following elements and functions:
[0143] 120A, B, C, D: Audio device playback sounds corresponding to rendered content being played back by audio devices 100A-100D within the same acoustic space.
[0144] 204A, B, C, D: A group of DSSS signal replicas corresponding to the group of DSSS signals generated by other audio devices in the audio environment (in this example, at least audio devices 100B, 100C, and 100D). In this example, the DSSS signal replicas 204A to 204D are provided by the orchestration module 213A. Here, the orchestration module 213A provides the DSSS information 204B to 204D to the audio devices 100B to 100D, for example, via wireless communication.
[0145] 205A, B, C, D: These elements correspond to DSSS information belonging to and / or used by each of the audio devices 100A to 100D. DSSS information 205A may include parameters (one or more DSSS spreading code parameters, one or more DSSS carrier parameters, etc.) used by the control system 160 of audio device 100A for purposes such as generating, modulating, or demodulating DSSS signal sets. DSSS information 205B, 205C, and 205D may include parameters (e.g., one or more DSSS spreading code parameters and one or more DSSS carrier parameters) used by audio devices 100B, 100C, and 100D, respectively, for purposes such as generating, modulating, or demodulating DSSS signal sets. In some examples, DSSS information 205A to 205D may indicate which audio device is generating the acoustic DSSS signal sets.
[0146] 213A: Orchestration module. In this example, the orchestration module 213A generates DSSS information 205A to 205D and provides DSSS information 205A to the DSSS signal generator 212A, DSSS information 205A to 205D to the DSSS signal demodulator, and DSSS information 205B to 205D to the audio devices 100B to 100D, for example via wireless communication. In some examples, the orchestration module 213A generates DSSS information 205A to 205D based at least in part on information 235A to 235D and / or acoustic scene metric group 225A to 225D.
[0147] 214A: A DSSS signal demodulator configured to demodulate a group of microphone signals 206A received by at least one or more microphones 111A. In this example, the DSSS signal demodulator 214A outputs a group of demodulated coherent baseband signals 208A. In several alternative forms, the DSSS signal demodulator 214A can receive and demodulate microphone signal groups 206B-206D from audio devices 100B-100D and output a group of demodulated coherent baseband signals 208B-208D.
[0148] 218A: A baseband processor configured for baseband processing of at least a group of demodulated coherent baseband signals 208A, and in some examples, the group of demodulated coherent baseband signals 208B-208D received from audio devices 100B-100D. In this example, the baseband processor 218A is configured to output one or more estimated acoustic scene metrics 225A-225D. In some embodiments, the baseband processor 218A is configured to determine the group of acoustic scene metrics 225B-225D based on the group of demodulated coherent baseband signals 208B-208D received from audio devices 100B-100D. However, in some cases, the baseband processor 218A (or the acoustic scene metric processing module 233A) may receive the group of acoustic scene metrics 225B-225D from audio devices 100B-100D.
[0149] 233A: An acoustic scene metric processing module configured to receive and apply a group of acoustic scene metrics 225A to 225D. In this example, the acoustic scene metric processing module 233A is configured to generate information 235A to 235D at least partially based on the group of acoustic scene metrics 225A to 225D and / or at least one audio device characteristic. The audio device characteristic(s) may correspond to one or more audio devices 100A and / or audio devices 100B to 100D.
[0150] Figure 4 is a block diagram showing an example of audio device elements according to another disclosed embodiment. As with other figures provided herein, the types and number of elements shown in Figure 4 are given merely as examples. Other embodiments may include more, fewer, and / or different types and numbers of elements. In this example, the audio device 100A in Figure 4 is an instance of the device 150 described above with reference to Figures 1B, 2, and 3. The example in Figure 4 includes all the elements in Figure 3 plus additional elements. Elements common to Figures 2 and 3 are not described again here except that their function may differ in the embodiment of Figure 4.
[0151] In this configuration, the control system 160 is configured to process the received microphone signal group 206A to generate a pre-processed microphone signal group 207A. In some configurations, processing the received microphone signal group may include applying a bandpass filter and / or echo cancellation. In this example, the control system 160 (more specifically the DSSS signal demodulator 214A) is configured to extract the DSSS signal group from the pre-processed microphone signal group 207A.
[0152] In this example, the microphone system 111A includes an array of microphones, which in some cases is or may include one or more directional microphones. In this configuration, processing the received microphone signal group includes receiving beamforming via a beamformer 215A in this example. In this example, the pre-processed microphone signal group 207A output by the beamformer 215A is or includes a group of spatial microphone signals.
[0153] In this configuration, the DSSS signal demodulator 214A processes the spatial microphone signal group, thereby improving performance for audio systems where audio devices are spatially distributed throughout the audio environment. Receiver beamforming is one way to avoid the aforementioned "near-far problem." For example, the control system 160 may be configured to use beamforming to compensate for closer and / or louder audio devices to receive audio device playback from farther and / or quieter audio devices.
[0154] Receiver beamforming may include, for example, delaying and multiplying the signals from each microphone in a microphone array by different coefficients. In some examples, the beamformer 215A may apply a Dorkh Chebyshev weighting pattern. However, in other embodiments, the beamformer 215A may apply different weighting patterns. According to some such examples, a main lobe may be generated along with a null and side lobes. In addition to controlling the main lobe width (beam width) and side lobe levels, in some examples, the position of the null can also be controlled.
[0155] Signals that are not audible In some cases, the DSSS signal components of audio device playback may be inaudible to people in the audio environment. In some such cases, the content stream components of audio device playback may cause perceptual masking of the DSSS signal components of audio device playback.
[0156] Figure 5 is a graph showing an example of the levels of the content stream component and the DSSS signal component of the audio playback sound from an audio device, over a certain frequency range. In this example, curve 501 corresponds to the level of the content stream component, and curve 530 corresponds to the level of the DSSS signal component.
[0157] A DSSS signal typically includes data, a carrier signal, and a spreading code. If we omit the need to transmit the data over the channel, the modulated signal s(t) can be expressed as follows:
[0158]
number
[0159] In the above equation, A represents the amplitude of the DSSS signal, C(t) represents the spreading code, and Sin() represents the sinusoidal carrier at carrier frequency f0Hz. Curve 530 in Figure 5 corresponds to an example of s(t) in the above equation.
[0160] One of the potential advantages of several disclosed modes related to acoustic DSSS signal sets is that by spreading the signal, the amplitude of the DSSS signal component is reduced relative to a given amount of energy in the acoustic DSSS signal, thereby reducing the perceptibility of the DSSS signal component in the sound reproduced by the audio device.
[0161] This allows the DSSS signal component of the audio device's playback sound (e.g., represented by curve 530 in Figure 5) to be placed at a level sufficiently lower than the level of the content stream component of the audio device's playback sound (e.g., represented by curve 501 in Figure 5) so that the DSSS signal component is not perceived by the listener. Some disclosed modes of operation involve optimizing the parameters of the DSSS signal to maximize the signal-to-noise ratio (SNR) of the derived set of DSSS signal observations and / or reduce the perceived probability of the DSSS signal component, by utilizing the masking properties of the human auditory system. Some disclosed examples include applying weights to the levels of the content stream component and / or the levels of the DSSS signal component. Some such examples involve applying noise compensation methods in which the acoustic DSSS signal component is treated as a signal and the content stream component is treated as noise. Some such examples include applying one or more weights according to (e.g., proportionally to) a playback / listening target metric.
[0162] DSSS spreading code As described elsewhere in this specification, in some examples the DSSS information 205 provided by the orchestration device (for example, provided by the orchestration module 213A described above with reference to Figure 3) may include one or more DSSS diffusion code parameters.
[0163] The spreading code used to spread the carrier wave to generate DSSS signals (one or more) is extremely important. The set of DSSS spreading codes is preferably selected such that the corresponding DSSS signal set has the following characteristics:
[0164] 1. Sharp principal lobes in the autocorrelation waveform. 2. Low side lobes in non-zero delay autocorrelation waveforms. 3. When multiple devices access the media simultaneously (for example, when simultaneously playing a set of modified audio playback signals containing DSSS signal components), the cross-correlation between any two spreading codes in the spreading code set used should be low. 4. The DSSS signal is unbiased (the DC component is zero).
[0165] A family of spreading codes (e.g., gold codes commonly used in GPS contexts) typically features the four points described above. When multiple audio devices are all simultaneously reproducing a modified audio playback signal set containing DSSS signal components, and each audio device uses a different spreading code (e.g., low cross-correlation, while having good cross-correlation characteristics), a receiving audio device should be able to simultaneously receive and process all acoustic DSSS signals by using a code-domain multiple access (CDMA) scheme. Using a CDMA scheme, multiple audio devices can simultaneously transmit acoustic DSSS signals, sometimes using a single frequency band. The spreading code is generated at runtime and / or pre-generated and stored in memory in a data structure such as a lookup table.
[0166] To implement DSSS, two-phase-shifted modulation (BPSK) modulation can be used in some examples. Furthermore, in some examples, DSSS spreading codes may be interplexed orthogonally to each other, for example, as follows, to implement a four-phase-shifted modulation (QPSK) system:
[0167]
number
[0168] In the above formula, A I and A Q These represent the amplitudes of the in-phase signal and the quadrature signal, respectively, and C I and C Q∫ represents the code sequences of the in-phase and quadrature signals, respectively, and f0 represents the center frequency (8200) of the DSSS signal. The above are examples of coefficients for parameterizing the DSSS carrier and DSSS spreading code, based on several examples. These parameters are examples of the DSSS information 205 described above. As described above, the DSSS information 205 may be provided by an orchestration device such as the orchestration module 213A, and may be used, for example, by the signal generator block 212 to generate the DSSS signal.
[0169] Figure 6 is a graph showing an example of the power of two DSSS signals with different bandwidths but the same center frequency. In these examples, Figure 6 shows the spectra of two DSSS signals 630A and 630B centered at the same center frequency 605. In some examples, DSSS signal 630A may be generated by one audio device in the audio environment (e.g., audio device 100A), and DSSS signal 630B may be generated by another audio device in the audio environment (e.g., audio device 100B).
[0170] In this example, DSSS signal 630B is chipped at a higher rate than DSSS signal 630A (in other words, more bits per second are used in the spread spectrum), and as a result, the bandwidth of DSSS signal 630B is greater than the bandwidth of DSSS signal 630A. For a given amount of energy in each DSSS signal, the larger the bandwidth of DSSS signal 630B, the lower the amplitude and perceptibility of DSSS signal 630B will be compared to DSSS signal 630A. Higher bandwidth DSSS signals also result in higher delay resolution of the baseband data product, leading to higher resolution estimates of acoustic scene metrics based on DSSS signals (such as time-of-flight estimates, time-of-arrival (ToA) estimates, range estimates, and direction-of-arrival (DoA) estimates). However, higher bandwidth DSSS signals also widen the receiver noise bandwidth, resulting in a lower SNR of the extracted acoustic scene metrics. Furthermore, if the bandwidth of the DSSS signal is too large, coherence and fading problems associated with the DSSS signal may occur.
[0171] The amount of cross-correlation rejection is limited by the length of the spreading code used to generate the DSSS signal. For example, a 10-bit gold code has only -26 dB of rejection for adjacent codes. This can lead to an example of the near-far problem described above, where a relatively low-amplitude signal is obscured by the cross-correlation noise of another, louder signal. Part of the novelty of the systems and methods described herein includes orchestration schemes designed to mitigate or avoid such problems.
[0172] Orchestration methods Figure 7 shows elements of an orchestration module in one example. As with other figures provided herein, the types and number of elements shown in Figure 7 are given merely as examples. In other forms, more, fewer, and / or different types and numbers of elements may be included. According to some examples, the orchestration module 213 may be implemented by an instance of the apparatus 150 described above with reference to Figure 1B. In some such examples, the orchestration module 213 may be implemented by an instance of the control system 160. In some examples, the orchestration module 213 may be an instance of the orchestration module described above with reference to Figure 3. In some such examples,
[0173] In this configuration, the orchestration module 213 includes a perceptual model application module 710, an acoustic model application module 711, and an optimization module 712.
[0174] In this example, the perceptual model application module 710 is configured to apply a model of the human auditory system to make one or more perceptual effect estimates 702 of the perceptual effect of a set of acoustic DSSS signals on a listener in an acoustic space, at least in part on priori information 701. The acoustic space may be, for example, an audio environment in which a set of audio devices organized by the orchestration module 213 is placed, a room in such an audio environment, and so on. The estimate(s) 702 may change over time. In some examples, the perceptual effect estimate(s) 702 may be an estimate of the listener's ability to perceive the set of acoustic DSSS signals, for example, based on the type and level of audio content (if any) currently being played in the acoustic space. The perceptual model application module 710 may be configured to apply one or more models of auditory masking, for example, masking as a function of frequency and loudness, spatial auditory masking, and so on. The perceptual model application module 710 may be configured to apply one or more models of human loudness perception, such as human loudness perception as a function of frequency.
[0175] In some examples, the a priori information 701 may be, or include, information relating to the acoustic space, information relating to the transmission of acoustic DSSS signals in the acoustic space, and / or information relating to a listener who is known to be using the acoustic space. For example, the a priori information 701 may include the number of audio devices in the acoustic space (e.g., audio devices being orchestrated), the location of the audio devices, the capabilities of the loudspeaker system and / or microphone system of the audio devices, information about the impulse response of the audio environment, information about one or more doors and / or windows of the audio environment, and information about the audio content currently being played in the acoustic space. In some examples, the a priori information 701 may include information about the auditory capabilities of one or more listeners.
[0176] In this configuration, the acoustic model application module 711 is configured to perform one or more acoustic DSSS signal performance estimations 703 of an acoustic DSSS signal group in an acoustic space, at least in part, based on prior information 701. For example, the acoustic model application module 711 may be configured to estimate the extent to which the microphone system of each audio device can detect an acoustic DSSS signal group from other audio devices in the acoustic space, which may be referred to herein as a form of "mutual audibility" of audio devices. Such mutual audibility may, in some cases, be an acoustic scene metric previously estimated by a baseband processor, at least in part, based on a previously received acoustic DSSS signal group. In some such configurations, the mutual audibility estimation may be part of the prior information 701, and in some such configurations, the orchestration module 213 may not include the acoustic model application module 711. However, in some configurations, the mutual audibility estimation may be performed independently by the acoustic model application module 711.
[0177] In this example, the optimization module 712 is configured to determine the DSSS parameters 705 of all audio devices organized by the orchestration module 213, at least in part, based on perceptual effect estimates (single or multiple) 702, a set of acoustic DSSS signal performance estimates 703, and current playback / listening purpose information 704. The current playback / listening purpose information 704 may, for example, indicate the relative need for a new set of acoustic scene metrics based on the acoustic DSSS signal set.
[0178] For example, when one or more audio devices are newly powered on in an acoustic space, there may be a high need for a new set of acoustic scene metrics related to audio device auto-positioning, audio device inter-audibility, etc. At least a portion of the new set of acoustic scene metrics can be based on the acoustic DSSS signal set. Similarly, when existing audio devices are moved within an acoustic space, there may be a high need for a new set of acoustic scene metrics. Likewise, when a new noise source is present in or near the acoustic space, there may be a high need to determine a new set of acoustic scene metrics.
[0179] If the current playback / listening purpose information 704 indicates a high need to determine a new set of acoustic scene metrics, the optimization module 712 may be configured to determine the DSSS parameter set 705 by giving relatively higher weight to acoustic DSSS signal performance estimates (single or multiple) 703 than to perceptual effect estimates (single or multiple) 702. For example, the optimization module 712 may be configured to determine the DSSS parameter set 705 by emphasizing the system's ability to generate a high SNR observation set of acoustic DSSS signals, and not emphasizing the impact / perceptibility of the acoustic DSSS signals to the user. In some such examples, the DSSS parameter set 705 may correspond to an audible acoustic DSSS signal.
[0180] However, if no recent changes have been detected in or near the acoustic space, and at least one or more acoustic scene metrics have been initially estimated, the need for a new set of acoustic scene metrics may not be high. If no recent changes have been detected in or near the acoustic space, at least one or more acoustic scene metrics have been initially estimated, and the audio content is currently being played in the acoustic space, the relative importance of immediately estimating one or more new acoustic scene metrics may be even lower.
[0181] If the current playback / listening purpose information 704 indicates a low level of need to determine a new set of acoustic scene metrics, the optimization module 712 may be configured to determine the DSSS parameter set 705 by giving relatively lower weight to the acoustic DSSS signal performance estimates 703 than to the perceptual effect estimates 702. In such examples, the optimization module 712 may be configured to determine the DSSS parameter set 705 by emphasizing the impact and perceptibility of the acoustic DSSS signal set to the user, rather than emphasizing the system's ability to generate a set of high SNR observations of the acoustic DSSS signal set. In some such examples, the DSSS parameter set 705 may correspond to a sub-audible acoustic DSSS signal set.
[0182] As will be discussed later in this book (for example, in other examples of audio device orchestration), the parameters of the acoustic DSSS signal set provide a rich variety of ways in which the orchestration device can modify the acoustic DSSS signal set to enhance the performance of the audio system.
[0183] Figure 8 shows another example of an audio environment. In Figure 8, audio devices 100B and 100C are located at distances 810 and 811, respectively, from device 100A. In this particular situation, distance 811 is greater than distance 810. Assuming that audio devices 100B and 100C are producing audio device playback sounds at approximately the same level, this means that audio device 100A receives the acoustic DSSS signal set from audio device 100C at a lower level than the acoustic DSSS signal set from audio device 100B, due to additional acoustic loss due to the longer distance of 811. In some embodiments, audio devices 100B and 100C can be organized to enhance audio device 100A's ability to extract the acoustic DSSS signal set and determine the acoustic scene metric set based on the acoustic DSSS signal set.
[0184] Figure 9 shows an example of the main lobes of the acoustic DSSS signal group generated by audio devices 100B and 100C in Figure 8. In this example, these acoustic DSSS signals have the same bandwidth and are located at the same frequency, but have different amplitudes. Here, the main lobe of acoustic DSSS signal 230B is generated by audio device 100B, and the main lobe of acoustic DSSS signal 230C is generated by audio device 100C. According to this example, the peak power of acoustic DSSS signal 230B is 905B, and the peak power of acoustic DSSS signal 230C is 905C. Here, acoustic DSSS signal 230B and acoustic DSSS signal 230C have the same center frequency 901.
[0185] In this example, the orchestration device (which in some cases includes an instance of the orchestration module 213 in Figure 7, and in some cases may be the audio device 100A in Figure 8) enhances the ability of audio device 100A to extract the acoustic DSSS signal group by equalizing the digital levels of the acoustic DSSS signals generated by audio devices 100B and 100C, so that the peak power of acoustic DSSS signal 230C is greater than the peak power of acoustic DSSS signal 230B by a coefficient that offsets the difference in acoustic loss due to the difference in distances 810 and 811. Thus, according to this example, audio device 100A receives the acoustic DSSS signal group 230B from audio device 100C at approximately the same level as the acoustic DSSS signal group received from audio device 100B, due to the additional acoustic loss due to the longer distance 811.
[0186] The surface area around a point sound source increases with the square of the distance from the source. This means that, according to the inverse square law, the same sound energy from the source is distributed over a wider area, and the energy intensity decreases with the square of the distance from the source. If we set the distance 810 to b and the distance 811 to c, the sound energy that audio device 100A receives from audio device 100B is 1 / b 2is proportional to, and the sound energy received by the audio device 100A from the audio device 100C is 1 / c 2 is proportional to. The difference in sound energy is 1 / (c 2 -b 2 ) is proportional to. Thus, in some embodiments, the orchestration device can multiply the energy generated by the audio device 100C by (c 2 -b 2 ). This is an example of a method of changing the DSSS parameter set for performance improvement.
[0187] In some embodiments, the optimization process may be more complex and may consider more factors than the inverse square law. In some examples, equalization may be performed via the full-band gain applied to the DSSS signal or via an equalization (EQ) curve that enables equalization of the non-flat (frequency-dependent) response of the microphone system 110A.
[0188] Figure 10 is a graph illustrating an example of a time-domain multiple access (TDMA) scheme. One way to avoid the near-far problem is to organize multiple audio devices that transmit and receive acoustic DSSS signals so that each audio device has a different time slot scheduled for reproducing its acoustic DSSS signals. This is known as the TDMA scheme. In the example shown in Figure 10, an orchestration device causes audio devices 1, 2, and 3 to emit acoustic DSSS signals according to the TDMA scheme. In this example, audio devices 1, 2, and 3 radiate acoustic DSSS signals in the same frequency band. According to this example, the orchestration device causes audio device 3 to radiate acoustic DSSS signals from time t0 to time t1, then the orchestration device causes audio device 2 to radiate acoustic DSSS signals from time t1 to time t2, then the orchestration device causes audio device 1 to radiate acoustic DSSS signals from time t2 to time t3, and so on.
[0189] Therefore, in this example, the two DSSS signals are never transmitted or received simultaneously. Consequently, the remaining DSSS signal parameters, such as amplitude, bandwidth, and length (as long as each DSSS signal remains within its allocated time slot), are irrelevant to multiple access. However, these DSSS signal parameters do affect the quality of the observations extracted from the DSSS signals.
[0190] Figure 11 is a graph illustrating an example of a frequency domain multiple access (FDMA) scheme. In some cases (for example, due to limited bandwidth of the DSSS signal set), an orchestration device may be configured to allow an audio device to simultaneously receive the acoustic DSSS signal set from two other audio devices in the audio environment. In some such cases, if each audio device transmitting the acoustic DSSS signal set reproduces its respective acoustic DSSS signal set in a different frequency band, the received power levels of the acoustic DSSS signal set will differ significantly. This is the FDMA scheme. In the example of the FDMA scheme shown in Figure 11, the main lobes of the DSSS signal sets 230B and 230C are transmitted simultaneously from different audio devices, but their center frequencies are different (f1, f2) and their frequency bands are also different (b1, b2). In this example, the frequency bands b1 and b2 of the main lobes do not overlap. Such an FDMA scheme is advantageous when the acoustic losses associated with the paths of the acoustic DSSS signal set differ significantly.
[0191] In some embodiments, the orchestration device may be configured to vary the FDMA, TDMA, or CDMA scheme to mitigate near-far problems. In some examples, the length of the DSSS spreading code may be varied depending on the relative audibility of the devices in the room. As described above with reference to Figure 6, if the amount of energy of the acoustic DSSS signal is the same, increasing the bandwidth of the acoustic DSSS signal by spreading the code results in a relatively lower maximum power and relatively reduced audibility of the acoustic DSSS signal. Alternatively or additionally, in some embodiments, the DSSS signal groups may be arranged orthogonally to each other. Such an implementation allows the system to have DSSS signal groups with different spreading code lengths simultaneously. Alternatively or additionally, in some embodiments, the energy of each DSSS signal may be modified to reduce the effects of near-far problems (e.g., to increase the level of acoustic DSSS signals produced by relatively low-volume and / or more distant transmitting audio devices) and / or to obtain an optimal signal-to-noise ratio for a given operating purpose.
[0192] Figure 12 is a graph illustrating other examples of orchestration methods. The elements of Figure 12 are as follows:
[0193] 1210, 1211, and 1212: Frequency bands that do not overlap with each other. 230Ai, Bi, and Ci: Multiple acoustic DSSS signals time-division multiplexed within frequency band 1210. Although audio devices 1, 2, and 3 may appear to be using different portions of frequency band 1210, in this example, the main lobes of the acoustic DSSS signal group 230Ai, Bi, and Ci extend across most or all of frequency band 1210. 230D and E: Multiple acoustic DSSS signals code-division multiplexed within frequency band 1211. Although audio devices 4 and 5 may appear to be using different portions of frequency band 1211, in this example, the main lobes of the acoustic DSSS signal groups 230D and 230E extend across most or all of frequency band 1211. 230Aii, Bii, and Cii: Multiple acoustic DSSS signals code-division multiplexed within frequency band 1212. Although audio devices 1, 2, and 3 may appear to use different portions of frequency band 1210, in this example, the main lobes of the acoustic DSSS signal group 230Aii, Bii, and Cii extend across most or all of frequency band 1212.
[0194] Figure 12 illustrates an example of how TDMA, FDMA, and CDMA may be used in combination in a particular embodiment of the present invention. In frequency band 1 (1210), TDMA is used to organize the acoustic DSSS signal groups 230Ai, Bi, and Ci transmitted by audio devices 1-3, respectively. Frequency band 1210 is a single frequency band such that the acoustic DSSS signal groups 230Ai, Bi, and Ci cannot simultaneously reside within it without overlap.
[0195] In frequency band 2(1211), CDMA is used to organize acoustic DSSS signal groups 230D and E from audio devices 4 and 5, respectively. In this particular example, acoustic DSSS signal 230D is generated by using a longer DSSS spreading code than the one used to generate acoustic DSSS signal 230E. A shorter DSSS spreading code duration is useful for audio device 5 when it is louder than audio device 4, as seen from the receiving audio device, because a shorter DSSS spreading code duration results in a wider bandwidth and a lower peak frequency for the resulting DSSS signal. The signal-to-noise ratio (SNR) may also be improved by the relatively longer DSSS spreading code duration of acoustic DSSS signal 230D.
[0196] In frequency band 3 (1212), CDMA is used to organize the acoustic DSSS signal groups 230Aii, Bii, and Cii transmitted by audio devices 1-3, respectively. These acoustic DSSS signal groups are alternative codes transmitted by audio devices 1-3, which are simultaneously transmitting TDMA-organized acoustic DSSS signal groups to the same audio devices in frequency band 1210. This is a form of FDMA in which longer spreading codes are placed and transmitted simultaneously within one frequency band (1212) (without TDMA), while shorter spreading codes are placed in another frequency band (1210) where TDMA is used.
[0197] Figure 13 is a graph illustrating another example of the orchestration method. In this configuration, audio device 4 transmits mutually orthogonal acoustic DSSS signal sets 230Di and 230Dii, and audio device 5 transmits mutually orthogonal acoustic DSSS signal sets 230Ei and 230Eii. In this example, all acoustic DSSS signal sets are transmitted simultaneously within a single frequency band 1310. In this case, the orthogonal acoustic DSSS signal sets 230Di and 230Ei are longer than the in-phase signals 230Dii and 230Eii transmitted by the two audio devices. As a result, each audio device has a higher set of SNR observations derived from acoustic DSSS signal sets 230Di and 230Ei, as well as a faster, noisier set of observations derived from acoustic DSSS signal sets 230Dii and 230Eii, albeit at a lower update rate. This is an example of a CDMA-based orchestration method in which two audio devices transmit acoustic DSSS signal sets designed for an acoustic space shared by the two audio devices. In some cases, the orchestration method may be based at least partially on the current listening purpose.
[0198] Figure 14 shows elements of an audio environment in another example. In this example, the audio environment 1401 is a multi-room dwelling including acoustic spaces 130A, 130B, and 130C. In this example, doors 1400A and 1400B can alter the coupling of each acoustic space. For example, when door 1400A is open, acoustic spaces 130A and 130C are acoustically coupled at least to some extent, whereas when door 1400A is closed, acoustic spaces 130A and 130C are not acoustically coupled to a significant degree. In some embodiments, the orchestration device may be configured to detect whether a door is open (or whether another acoustic obstruction has been moved) in response to the detection or absence of audio device playback sound in an adjacent acoustic space.
[0199] In some examples, the orchestration device can organize all of the audio devices 100A-100E in all of the acoustic spaces 130A, 130B, and 130C. However, due to a significant level of acoustic separation between acoustic spaces 130A, 130B, and 130C when doors 1400A and 1400B are closed, in some examples, the orchestration device can treat acoustic spaces 130A, 130B, and 130C as independent when doors 1400A and 1400B are closed. In some examples, the orchestration device can treat acoustic spaces 130A, 130B, and 130C as independent even when doors 1400A and 1400B are open. However, in some cases, the orchestration device may manage audio devices located near door 1400A and / or 1400B such that when the acoustic spaces are combined by the opening of the door, audio devices near the open door are treated as corresponding to the rooms on either side of the door. For example, if the orchestration device determines that door 1400A is open, the orchestration device may be configured to consider audio device 100C as an audio device in acoustic space 130A and also as an audio device in acoustic space 130C.
[0200] Figure 15 is a flowchart outlining another example of the method for orchestrating audio devices disclosed herein. The blocks of Method 1500, as with other methods described herein, are not necessarily performed in the order shown. Also, such a method may include more or fewer blocks than those illustrated and / or described. Method 1500 may be performed by a system including an orchestration device and an audio device to be orchestrated. The system may include instances of the apparatus 150 shown in Figure 1B and described above, one of which is configured as the orchestration device. In some examples, the orchestration device may include an instance of the orchestration module 213 disclosed herein.
[0201] In this example, block 1505 includes the steady-state operation of all participating audio devices. In this context, “steady-state” operation means operating according to the set of parameters most recently received from the orchestration device. In this mode, the set of parameters includes one or more DSSS spreading code parameters and one or more DSSS carrier parameters.
[0202] In this example, block 1505 also includes one or more devices waiting for a trigger condition. The trigger condition may be, for example, an acoustic change in the audio environment in which the audio device to be orchestrated is located. The acoustic change may be, or include, noise from a noise source, a change corresponding to the opening or closing of a door or window (e.g., an increase or decrease in the audibility of sound played from one or more loudspeakers in an adjacent room), detected movement of an audio device in the audio environment, detected movement of a person in the audio environment, detected speech of a person in the audio environment (e.g., a wake word), the start of playback of audio content (e.g., the start of movie, television program, music content, etc.), or a change in the playback of audio content (e.g., a volume change exceeding a threshold change in decibels). In some examples, the acoustic change is detected via, for example, a set of acoustic DSSS signals as disclosed herein (e.g., one or more acoustic scene metrics 225A estimated by the baseband processor 218 of the audio device in the audio environment).
[0203] In some examples, the trigger condition may indicate that a new audio device has been powered on in an audio environment. In some such examples, the new audio device may be configured to produce one or more characteristic sounds, which may or may not be audible to humans. According to some examples, the new audio device may be configured to reproduce an acoustic DSSS signal according to a type of DSSS spreading code reserved for the new device. Some examples of reserved DSSS spreading codes are described below.
[0204] In this example, block 1510 determines whether the trigger condition has been detected. If so, the process proceeds to block 1515. Otherwise, the process returns to block 1505. In some forms, block 1505 may include block 1510.
[0205] In this example, block 1515 includes the orchestration device determining one or more updated sets of acoustic DSSS parameters for one or more (possibly all) of the audio devices being orchestrated, and providing the updated acoustic DSSS parameters(s) to the audio devices(s) being orchestrated. In some examples, block 1515 may include the orchestration device providing DSSS information 205 as described elsewhere in this specification. The determination of the updated acoustic DSSS parameters(s) may include using existing knowledge and estimates of acoustic spaces, such as:
[0206] • Device location • Device range; • Device orientation and relative angle of incidence; • Relative clock bias and skew between devices; • Relative audibility of the device; • Estimated indoor noise level; • The number of microphones and loudspeakers in each device; • Directivity of the loudspeakers in each device; • Microphone directivity of each device; • The type of content being rendered in the acoustic space; • The position of one or more listeners in the acoustic space; and / or Knowledge of acoustic spaces, including specular reflection and occlusion.
[0207] These factors may, in some cases, be combined with a specific operational objective to determine new operational points. It should be noted that many of these parameters, used as existing knowledge when determining the updated DSSS parameter set, can themselves be derived from the acoustic DSSS parameter set. Therefore, in some cases, it is easy to see that the performance of an organized acoustic DSSS system can be iteratively improved as the system acquires more information, more accurate information, and so on.
[0208] In this example, block 1520 includes one or more parameters used to generate a set of acoustic DSSS signals, according to updated acoustic DSSS parameters (one or more) received from the orchestration device, by one or more audio devices undergoing orchestration. In this configuration, after block 1520 is completed, the process returns to block 1505. Although termination is not shown in the flowchart of Figure 15, method 1500 can be terminated in various ways, for example, when the audio device is powered off.
[0209] Figure 16 shows another example of an audio environment. The audio environment 130 shown in Figure 16 is similar to that shown in Figure 8, but also shows the angular distance between audio devices 100B and 100C as seen from audio device 100A (relative to audio device 100A). In Figure 16, audio devices 100B and 100C are separated from device 100A by distances 810 and 811, respectively. In this particular situation, distance 811 is greater than distance 810. Assuming that audio devices 100B and 100C are producing audio device playback sound at approximately the same level, this means that audio device 100A receives the acoustic DSSS signals from audio device 100C at a lower level than the acoustic DSSS signals from audio device 100B, due to additional acoustic loss because the distance 811 is longer.
[0210] In this example, the focus is on the arrangement of devices 100B and 100C to optimize the ability of device 100A to hear both devices 100B and 100C. As outlined above, there are other factors to consider, but in this example, the focus is on the diversity of angles of arrival resulting from the angular separation of audio devices 100B and 100C relative to audio device 100A. Due to the difference in distances 810 and 811, the arrangement results in longer code lengths for audio devices 100B and 100C, mitigating the near-far problem by reducing cross-channel correlation. However, if the receiving beamformer (215) is implemented by audio device 100A, the angular separation between audio devices 100B and 100C places the microphone signal groups corresponding to the sounds from audio devices 100B and 100C on different lobes, providing further separation of the two received signals, thus mitigating the near-far problem to some extent. Therefore, this further separation allows the orchestration device to shorten the acoustic DSSS diffusion code length and obtain observation sets at a faster rate.
[0211] This is not limited to acoustic DSSS spread code length. If a spatial microphone feed is used by audio device 100A (and / or audio devices 100B and 100C) instead of an omnidirectional microphone feed, acoustic DSSS parameters that can be modified to mitigate near-far issues (even if FDMA or TDMA is used) may no longer be necessary.
[0212] Orchestration according to spatial means (in this case, angular diversity) depends on the availability of estimates of these characteristics. For example, the DSSS parameter set may be optimized for an omnidirectional microphone feed (206), and then, after the DoA estimate set becomes available, the acoustic DSSS parameter set may be optimized for a spatial microphone feed. This is one example of realizing the trigger condition described above, with reference to Figure 15.
[0213] Figure 17 is a block diagram showing examples of DSSS signal demodulator elements, baseband processor elements, and DSSS signal generator elements in several disclosed modes. As with other figures provided herein, the types and number of elements shown in Figure 17 are given merely as examples. Other modes may include more, fewer, and / or different types and numbers of elements. Other examples may implement other methods, such as frequency-domain correlation. In this example, the DSSS signal demodulator 214, baseband processor 218, and DSSS signal generator 212 are implemented by an instance of the control system 160 described above with reference to Figure 1B.
[0214] In some configurations, for each acoustic DSSS signal transmitted (reproduced) from each audio device that receives the acoustic DSSS signal group, there is one instance of the DSSS signal demodulator 214, baseband processor 218, and DSSS signal generator 212. In other words, in the configuration shown in Figure 16, audio device 100A implements one instance of the DSSS signal demodulator 214, baseband processor 218, and DSSS signal generator 212 corresponding to the acoustic DSSS signal group received from audio device 100B, and one instance of the DSSS signal demodulator 214, baseband processor 218, and DSSS signal generator 212 corresponding to the acoustic DSSS signal group received from audio device 100C.
[0215] For illustrative purposes, the following description of Figure 17 will continue to use this example of audio device 100A from Figure 16 as a local device implementing instances of the DSSS signal demodulator 214, baseband processor 218, and DSSS signal generator 212. More specifically, the following description of Figure 17 will assume that the microphone signal group 206 received by the DSSS signal demodulator 214 includes a reproduced sound containing acoustic DSSS signals generated by the loudspeaker of audio device 100B, and that the instances of the DSSS signal demodulator 214, baseband processor 218, and DSSS signal generator 212 shown in Figure 17 correspond to the acoustic DSSS signals reproduced by the loudspeaker of audio device 100B.
[0216] In this configuration, the DSSS signal generator 212 includes an acoustic DSSS carrier module 1715. The acoustic DSSS carrier module 1715 is configured to provide the DSSS signal demodulator 214 with a DSSS carrier replica 1705 of the DSSS carrier used by the audio device 100B to generate its acoustic DSSS signal set. In some alternative configurations, the acoustic DSSS carrier module 1715 may be configured to provide the DSSS signal demodulator 214 with one or more DSSS carrier parameters used by the audio device 100B to generate its acoustic DSSS signal set.
[0217] In this configuration, the DSSS signal generator 212 also includes an acoustic DSSS spreading code module 1720. The acoustic DSSS spreading code module 1720 is configured to provide the DSSS signal demodulator 214 with a DSSS spreading code 1706, which is used by the audio device 100B to generate its acoustic DSSS signal set. The DSSS spreading code 1706 corresponds to the spreading code C(t) in the formula disclosed herein. The DSSS spreading code 1706 may be, for example, a sequence of pseudorandom numbers (PRNs).
[0218] In this configuration, the DSSS signal demodulator 214 includes a bandpass filter 1703 configured to generate a bandpass filtered microphone signal group 1704 from the received microphone signal group 206. In some examples, the passband of the bandpass filter 1703 may be centered on the center frequency of the acoustic DSSS signal from the audio device 100B being processed by the DSSS signal demodulator 214. The bandpass filter 1703 may, for example, pass the main lobe of the acoustic DSSS signal. In some examples, the passband of the bandpass filter 1703 may be equal to the frequency band for transmitting the acoustic DSSS signal from the audio device 100B.
[0219] In this example, the DSSS signal demodulator 214 includes a multiplication block 1711A configured to convolve a bandpass-filtered microphone signal group 1704 with a DSSS carrier replica 1705 to generate a baseband signal group 1700. In this configuration, the DSSS signal demodulator 214 also includes a multiplication block 1711B configured to generate a non-spread baseband signal group 1701 by applying a DSSS spreading code 1706 to the baseband signal group 1700.
[0220] In this example, the DSSS signal demodulator 214 includes a accumulator 1710A, and the baseband processor 218 includes a accumulator 1710B. Accumulators 1710A and 1710B are also sometimes referred to as sum elements in this specification. Accumulator 1710A operates during a time that may be referred to herein as the “coherent time,” corresponding to the code length of each acoustic DSSS signal (in this example, the code length of the acoustic DSSS signal currently being played by the audio device 100B). In this example, accumulator 1710A performs an “integration and dumping” process. In other words, after taking the sum of the non-spread baseband signal group 1701 over the coherent time, accumulator 1710A outputs ("dumps") the demodulated coherent baseband signal 208 to the baseband processor 218. In some embodiments, the demodulated coherent baseband signal 208 may be a single number.
[0221] In this example, the baseband processor 218 includes a square law module 1712. In this example, the square law module is configured to square the absolute value of the demodulated coherent baseband signal 208 and output the power signal 1722 to the accumulator 1710B. After the absolute value and squaring process, the power signal can be considered a non-coherent signal. In this example, the accumulator 1710B operates over a non-coherent time. In some examples, the non-coherent time may be based on the input from the orchestration device. In some examples, the non-coherent time may be based on a desired SNR. According to this example, the accumulator 1710B outputs a delayed waveform 400 with multiple delays (also referred herein as "tau" or instances of tau (τ)).
[0222] The steps from 1704 to 208 in Figure 17 can be represented as follows:
number
[0223] In the above equation, Y(tau) represents the coherent demodulator output (208), d[n] represents the bandpass filtered signal (1704 or A in Figure 17), CA represents the local copy spreading the code used to modulate the DSSS signal by a distant device in the room (in this example, audio device 100B), and the last term is the carrier signal. In some examples, all of these signal parameters are orchestrated among audio devices in the audio environment (for example, determined and provided by an orchestration device).
[0224] From Y(tau)(208) in Figure 17<Y(tau)> The signal chain up to (400) is a non-coherent integral, and the coherent demodulator output is squared and averaged. The number of averagings (the number of times the non-coherent accumulator 1710B is performed) is a parameter that may be determined and provided by the orchestration device, for example, based on the determination that a sufficient SNR has been achieved in some examples. In some examples, an audio device implementing the baseband processor 218 may determine the number of averagings, for example, based on the determination that a sufficient SNR has been achieved.
[0225] Non-coherent integrals can be expressed mathematically as follows:
number
[0226] The aforementioned formula involves simply averaging the squared coherent delay waveform over a period defined by N, where N represents the number of blocks used in the non-coherent integral.
[0227] Figure 18 shows the elements of a DSSS signal demodulator in another example. In this example, the DSSS signal demodulator 214 is configured to generate a set of delay estimates, a set of DoA estimates, and a set of audible estimates. In this example, the DSSS signal demodulator 214 is configured to perform coherent demodulation, after which non-coherent integration is performed on the total delay waveform. Similar to the example described above with reference to Figure 17, this example assumes that the DSSS signal demodulator 214 is implemented by audio device 100A and configured to demodulate a set of acoustic DSSS signals to be reproduced by audio device 100B.
[0228] In this example, the DSSS signal demodulator 214 includes a bandpass filter 1703 configured to remove unwanted energy from a portion of the audio content being rendered for the listener's experience, as well as from other audio signals, such as a group of acoustic DSSS signals placed in other frequency bands to avoid near-far issues.
[0229] The matched filter 1811 is configured to calculate a delay waveform 1802 by correlating the bandpass-filtered signal 1704 with a local replica of the target acoustic DSSS signal. In this example, the local replica is an instance of the DSSS signal replica group 204 corresponding to the DSSS signal group generated by the audio device 100B. The matched filter output 1802 is then low-pass filtered by the low-pass filter 712 to generate a coherently demodulated complex delay waveform 208. In some alternative forms, the low-pass filter 712 may be placed after the squaring operation in the baseband processor 218 that generates a non-coherent averaged delay waveform, as in the example described above with reference to Figure 17.
[0230] In this example, the channel selector 1813 is configured to control the bandpass filter 1703 (e.g., the passband of the bandpass filter 1703) and the matched filter 1811 according to the DSSS information 205. As described above, the DSSS information 205 may include a set of parameters used by the control system 160 to demodulate the DSSS signal set. In some examples, the DSSS information 205 may indicate which audio device is generating the acoustic DSSS signal set. In some examples, the DSSS information 205 may be received from an external source (e.g., via wireless communication), such as an orchestration device.
[0231] Figure 19 is a block diagram showing examples of baseband processor elements in several disclosed modes. As with other figures provided herein, the types and number of elements shown in Figure 19 are given merely as examples. Other modes may include more, fewer, and / or different types and numbers of elements. In this example, the baseband processor 218 is implemented by an instance of the control system 160 described above with reference to Figure 1B.
[0232] In this particular mode, the coherent technique is not applied. Therefore, the first operation performed is to take the power of the complex delay waveform 208 via the square law module 1712 to generate a non-coherent delay waveform 1922. The non-coherent delay waveform 1922 is integrated by the accumulator 1710B for a certain period (specified in this example by the DSSS information 205 received from the orchestration device, but may be determined locally in some examples) to generate a non-coherent averaged delay waveform 400. According to this example, the delay waveform 400 is then processed in several ways as follows:
[0233] 1. The leading edge estimator 1912 is configured to perform a delay estimate 1902, which is an estimated time delay of the received signal. In some examples, the delay estimate 1902 may be based at least partially on an estimate of the leading edge position of the delayed waveform 400. According to some such examples, the delay estimate 1902 may be determined according to the number of time samples of the signal portion (e.g., the positive portion) of the delayed waveform (including itself), from the leading edge position of the delayed waveform 400 to a time sample of less than one chip period (inversely proportional to the signal bandwidth). In the latter case, this delay can be used to compensate for the width of the autocorrelation of the DSSS code. As the chipping rate increases, the width of the autocorrelation peak narrows and is minimized when the chipping rate is equal to the sampling rate. This condition (chipping rate equal to sampling rate) yields a delayed waveform 400 that best approximates the true impulse response of the audio environment to a given DSSS code. As the chipping rate increases, spectral overlap (aliasing) may occur following the DSSS signal modulator 220A. In some examples, if the chipping rate is equal to the sampling rate, the DSSS signal modulator 220A may be bypassed or omitted. A chipping rate that approaches that of the sampling rate (e.g., a chipping rate of 80% of the sampling rate, a chipping rate of 90% of the sampling rate, etc.) may provide a delayed waveform 400 that is a good approximation of the actual impulse response for some purposes. In some such examples, the delayed estimate 1902 may be based in part on information regarding the DSSS signal characteristics. In some examples, the leading edge estimater 1912 may be configured to estimate the position of the leading edge of the delayed waveform 400 according to the first instance of a value greater than a threshold during a certain time window. Some examples are described later with reference to Figure 20. In other examples, the leading edge estimater 1912 may be configured to estimate the position of the leading edge of the delayed waveform 400 according to the position of a maximum value (e.g., a local maximum within a certain time window), which is an example of "peak picking".It should be noted that many other techniques can be used to estimate latency (e.g., peak picking).
[0234] 2. In this example, the baseband processor 218 is configured to perform a DoA estimation 1903 by windowing the delayed waveform 400 (windowing block 1913) before using the delayed sum DoA estimator 1914. The delayed sum DoA estimator 1914 can perform a DoA estimation based at least in part on the determination of the steered response power (SRP) of the delayed waveform 400. Thus, the delayed sum DoA estimator 1914 may also be referred to herein as the SRP module or delayed sum beamformer. Windowing helps to separate time intervals around the leading edge so that the resulting DoA estimation is based more on the signal than on the noise. In some examples, the window size can range from tens of milliseconds to hundreds of milliseconds, for example, from 10 milliseconds to 200 milliseconds. In some examples, the window size may be selected based on knowledge of the decay time of a typical room or knowledge of the decay time of the audio environment in question. In some examples, the window size may be adaptively updated over time. For example, some embodiments may involve determining a window size such that at least some portion of the window is occupied by the signal portion of the delayed waveform 400. Some such embodiments may involve estimating the noise power according to time samples occurring before the leading edge. Some such embodiments may involve selecting a window size such that at least a threshold percentage of the window is occupied by portions of the delayed waveform corresponding to at least a threshold signal level, e.g., at least 6 dB greater than the estimated noise power, at least 8 dB greater than the estimated noise power, at least 10 dB greater than the estimated noise power, and so on.
[0235] 3. In this example, the baseband processor 218 is configured to perform an audible estimation 1904 by estimating signal versus noise power using the SNR estimation block 1915. In this example, the SNR estimation block 1915 is configured to extract a signal power estimate 402 and a noise power estimate 401 from the delayed waveform 400. In some such examples, the SNR estimation block 1915 may be configured to determine the signal portion and the noise portion of the delayed waveform 400, as described later with reference to Figure 20. In some such examples, the SNR estimation block 1915 may be configured to determine the signal power estimate 402 and the noise power estimate 401 by averaging the signal portion and the noise portion over a selected set of time windows. In some such examples, the SNR estimation block 1915 may be configured to perform an SNR estimation according to the ratio of the signal power estimate 402 to the noise power estimate 401. In some examples, the baseband processor 218 may be configured to perform an audible estimation 1904 according to the SNR estimation. Given a given amount of noise power, the signal-to-noise ratio (SNR) is proportional to the audibility of an audio device. Therefore, in some embodiments, the SNR may be used directly as a proxy (e.g., a proportional value) for an estimate of the audibility of an actual audio device. In some embodiments, including calibrated microphone feed sets, this may involve measuring absolute audibility (e.g., in dBSPL) and converting the SNR to an absolute audibility estimate. In some such embodiments, the method for determining the absolute audibility estimate takes into account acoustic loss due to the distance between audio devices and the variability of room noise. In other embodiments, other techniques may be used to estimate signal power, noise power, and / or relative audibility from delayed waveforms.
[0236] Figure 20 shows an example of a delayed waveform. In this example, the delayed waveform 400 is output by an instance of the baseband processor 218. In this example, the vertical axis represents power, and the horizontal axis represents the pseudorange in meters. As described above, the baseband processor 218 is configured to extract delay information, sometimes referred to herein as τ, from the demodulated acoustic DSSS signal. The value of τ can be converted to a pseudorange measurement, sometimes referred to as ρ, as follows:
[0237]
number
[0238] In the aforementioned equation, c represents the speed of sound. In Figure 20, the delayed waveform 400 includes a noise portion 2001 (sometimes called the noise floor) and a signal portion 2002. Negative values in the pseudorange measurements (and the corresponding delayed waveforms) can be identified as noise. Since negative ranges (distances) are not physically meaningful, power corresponding to negative pseudoranges is assumed to be noise.
[0239] In this example, the signal portion 2002 of waveform 400 includes a leading edge 2003 and a trailing edge. The leading edge 2003 is a prominent feature of the delayed waveform 400 when the power of the signal portion 2002 is relatively strong. In some examples, the leading edge estimator 1912 in Figure 19 may be configured to estimate the position of the leading edge 2003 according to the first instance of a power value greater than a threshold during a certain time window. In some examples, the time window may start when τ (or ρ) is zero. In some examples, the window size may range from tens of milliseconds to hundreds of milliseconds, for example, from 10 milliseconds to 200 milliseconds. According to some embodiments, the threshold may be a pre-selected value, for example, -5dB, -4dB, -3dB, -2dB, etc. In some alternative examples, the threshold may be based on the power of at least a portion of the delayed waveform 400, for example, the average power of the noise portion.
[0240] However, as described above, in other examples, the leading edge estimator 1912 may be configured to estimate the position of the leading edge 2003 according to the position of the maximum value (e.g., the maximum value within a time window). In some examples, the time window may be selected as described above.
[0241] The SNR estimation block 1915 of FIG. 19, in some examples, may be configured to determine an average noise value corresponding to at least a part of the noise portion 2001 and an average or peak signal value corresponding to at least a part of the signal portion 2002. The SNR estimation block 1915 of FIG. 19 may be configured to estimate the SNR by dividing the average signal value by the average noise value in some such examples.
[0242] FIG. 21 shows an example of a block according to another aspect. This example includes a correlator bank implementation of the DSSS signal demodulator 214. In this context, the term "correlator bank" means that multiple instances of an acoustic DSSS signal group are correlated at different delays. According to this example, the bulk delay estimator 2110 is used to coarsely align the DSSS correlator bank (214) so that only a subset of all delay groups need to be calculated by the baseband processor 218. In this aspect, the DSSS correlator bank (214) generates a windowed demodulated coherent baseband signal 208, and the baseband processor 218 generates a windowed non-coherent averaged delay waveform 400.
[0243] In this embodiment, the bulk delay estimator 2110 uses a signal rendered by a distant device as a reference to estimate the bulk delay. In such an example, the bulk delay estimator 2110 implements a cross-correlator that correlates a reference signal (2102) reproduced by another audio device ("distant device") within the audio environment with the received microphone signal group 206, and is configured to estimate the bulk delay 2103. Generally, the estimated bulk delay 2103 is different for each audio device that receives the acoustic DSSS signal group.
[0244] Some alternative embodiments include estimating the bulk delay 2103 according to the filter tap information of an acoustic echo canceller that cancels the reference playback of the distant device. The filter shows peaks corresponding to the direct signal groups from other devices, thereby obtaining a rough alignment.
[0245] The bulk delay estimator 2110 can improve efficiency by restricting subsequent "downstream" calculations. For example, through windowing processing, the pseudo range can be restricted to a range from x to y meters, such as 1 to 4 meters, 0 to 4 meters, 1 to 5 meters, -1 to 4 meters, etc., rather than the range shown in FIG. 20.
[0246] FIG. 22 shows an example of a block according to yet another embodiment. This example includes a "matched filter" version of the DSSS signal demodulator 214, which can be configured as described above with reference to FIG. 18 in some cases. This example also includes an instance of the bulk delay estimator 2110, which in this embodiment provides the bulk delay estimate 2103 to the baseband processor 218.
[0247] In this example, the window is steering (centered) by an external bulk delay estimate 2103 with respect to the signal components of the delayed waveform 2204 extracted using the windowing block 1913. An additional windowing block 2213 is centered using the bulk delay estimate 2103 and an offset 2206 to window the delayed waveform 400 in the noise-only region of the delayed waveform. For example, the offset windowed delayed waveform 2205 may correspond to the noise portion 2001 in Figure 20.
[0248] In this example, the baseband processor 218 windowes the delayed waveform 400 before performing SRP via the delayed sum beamformer 1914, as described above with reference to Figure 19. However, in this example, the baseband processor 218 controls the windowing block 1913 based on the bulk delay estimate 2103. In this configuration, the windowing block 1913 provides the windowed delayed waveform 2204 to the leading edge estimator 1912, the delayed sum beamformer 1914, and the SNR estimation block 1915. Furthermore, in this example, the baseband processor 218 controls the windowing block 2213 based on the bulk delay estimate 2103.
[0249] In some embodiments, the delay estimate 1902, estimated using the leading edge estimator 1912, may, in some examples, be used to window the subsequent acoustic DSSS observations. In some such embodiments, the delay estimate 1902 can replace the bulk delay 2103 in Figures 21 and 22.
[0250] Figure 23 is a block diagram showing examples of audio device elements in several disclosed modes. As with other figures provided herein, the types and number of elements shown in Figure 23 are given for illustrative purposes only. Other modes may include more, fewer, and / or different types and numbers of elements. In this example, the audio device 100A in Figure 23 is an instance of the apparatus 150 described above with reference to Figures 1B and 2-4. The mode shown in Figure 23 includes all the elements of Figure 4, except that the beamformer 215A in Figure 4 is replaced in Figure 23 by a more generalized preprocessing module 221A. Elements common to Figures 4 and 23 are not described again here, except that their function may differ in the mode of Figure 23.
[0251] In this configuration, the pre-processing module 221A is configured to pre-process the received microphone signal group 206A to generate a pre-processed microphone signal group 207A. In some configurations, pre-processing the received microphone signal group may include the application of a bandpass filter and / or cancellation. In some examples, the microphone system 111A may include an array of microphones, in some cases this microphone array may be or include one or more directional microphones. In some such examples, pre-processing the received microphone signal group may include receiver-side beamforming via the pre-processing module 221A.
[0252] Generally, each audio device has its own internal clock, which often functions independently of the clocks implemented by other audio devices in the audio environment. Clock offset, or bias, refers to clocks that are shifted by a specific amount of time (for example, the clock of audio device A and the clock of audio device B). Clocks generally operate at slightly different speeds, which is known as clock skew. Clock skew changes the clock bias over time. This change in clock bias alters the estimated range or distance between devices. This phenomenon is known as "range walk."
[0253] In systems where clock skew is limited by network synchronization and / or clock skew is estimated (possibly by techniques enumerated in this disclosure), it may be advantageous to limit the coherent integration time of the receiving device to mitigate SNR loss due to range walk during the integration period. In some examples, this can be combined with range walk compensation techniques, for example, when the skew is not large on the coherent integration time scale but is large on the non-coherent integration time scale.
[0254] Figure 24 shows a block of another example of the configuration. As with other figures provided herein, the types and number of elements shown in Figure 23 are given for illustrative purposes only. Other configurations may include more, fewer, and / or different types and numbers of elements. For example, in some configurations, the baseband processor 218 may include additional elements such as those described above with reference to Figures 19 and 22.
[0255] In this embodiment, with reference to Figure 15, one method of monitoring one of the types of trigger conditions mentioned above (for triggering an update of the acoustic DSSS parameter set) is implemented as a block configured to detect a change in the relative clock skew of any two audio devices in the audio environment. Several detailed examples of calculating the relative clock skew of two audio devices are shown below. In some examples, the enhanced coefficients of the DSSS signal demodulator 214 and the baseband processor 218 may be at least partially based on the relative clock skew. Furthermore, a change in clock skew greater than a threshold amount may, in some examples, result in a change in the global operating configuration (e.g., CDMA, FDMA, TDMA assignment) of all participating audio devices and may be a trigger condition that in some cases triggers the flow from block 1510 to block 1515 in Figure 15.
[0256] In the example shown in Figure 24, the DSSS signal generator 212A receives the signal skew parameter group 2402 and provides the DSSS signal replica group 204, corresponding to the DSSS signal group generated by other audio devices in the audio environment, to the DSSS signal demodulator 214. In some examples, the DSSS signal generator 212A may receive the DSSS signal replica group 204 and the signal skew parameter group 2402 from an orchestration device.
[0257] In the example shown in Figure 24, the DSSS signal demodulator 214 is shown receiving the DSSS signal replica group 204, the microphone signal group 206, and coherent integral time information 2401. In this example, the square law module 1712 of the baseband processor 218 is configured to receive the demodulated coherent baseband signal group 208 from the DSSS signal demodulator 214, generate a non-coherent delay waveform 1922, and provide the non-coherent delay waveform 1922 to the delay walk compensator 2410. In this example, the delay walk compensator 2410 is configured to compensate for the delay walk between the receiving audio device and the audio device whose acoustic DSSS signal is currently being processed by the baseband processor 218. In this example, the delay walk compensator 2410 is configured to compensate for the delay walk according to the received delay rate estimate 2403 and output a non-coherent compensated power delay waveform 2405. The term "delay walk" refers to the effect of non-zero delay rate terms, such as how much a delay waveform shifts over a given period. This is caused by a mismatch in the physical clock frequencies of the transmitting and receiving devices. In this example, the delay rate estimate 2403 is the rate of change of the estimated delay over time. According to some examples, the delay rate estimate 2403 may be determined according to a stored instance of delay estimates determined over a certain period (e.g., several hours, several days, several weeks, etc.). If the estimated delay rate is large, when the delay waveforms are integrated (averaged) non-coherently, the shift of the instantaneous delay waveform (e.g., the shift of the demodulated coherent baseband signal group 208 in Figure 24) will blur the final non-coherent averaged signal (e.g., signal 400 in Figure 24). As an example of the effect of a "significant" delay rate, considering a -3dB shift in the peak power response due to errors induced by the delay rate, delay rates exceeding the delay rate limit, expressed as delay_rate_lim in the following equation, induce errors worse than -3dB. In the following equation, T_code represents the temporal length of the entire spreading code sequence.
[0258]
number
[0259] In some examples, the delayed walk compensator 2410 can shift the signal (1922) before averaging it using the delayed rate estimate 2403. In some such examples, this shift is equal to the amount of delayed walks occurring over the non-coherent integration period, but the shift is applied in the reverse direction to cancel out the delayed walks.
[0260] In some alternative forms, the coherent processing occurring in the DSSS signal demodulator 214 may be modified according to clock bias and / or clock skew information. In one such example, a clock bias estimate is used in the DSSS signal generator 212 to shift the phase of the replica signal code (1720), so that the delay in the delayed waveform is attributable solely to the physical distance between the audio devices. In some examples, a set of clock skew estimates may be used in the DSSS signal generator 212 to shift the frequency of the replica signal carrier (1715) so that the resulting coherent waveform (208) has no residual frequency components (in other words, no sine waves remain). This condition can occur when the replica signal generates carriers corresponding to the physical signals transmitted by the audio devices currently being evaluated / listened to. These carrier frequencies are slightly different due to the different clock frequencies.
[0261] Figure 25 shows another example of an audio environment. In this example, the elements of Figure 25 are as follows:
[0262] 100i,j,k: Distributed audio devices that receive multiple orchestrations; 2500: A signal transmitted from audio device i (100i) and received by audio device j (100j); 2501: Signal transmitted from the audio device i (100i) and received by the audio device i (100i); 2502: Signal transmitted from the audio device j (100j) and received by the audio device i (100i); 2503: Signal transmitted from the audio device j (100j) and received by the audio device j (100j); 2510: Actual distance between the audio device i (100i) and the audio device j (100j), and 2511(i,j): Distance between the loudspeaker and the microphone of the audio device.
[0263] Next, some examples of asynchronous two-way ranging will be described with reference to FIG. 25. In this example, the audio devices are asynchronous and there is a bias between the clocks. In this particular mode, two-way ranging is used so that all unknown clock terms are canceled out. This particular example is executed with a pair of audio devices and is described with reference to audio devices 100i and 100j. A set of ranges between all audio devices in the acoustic space can be obtained by repeating this for all pairs of audio devices (e.g., audio device pairs 100i to 100k and audio device pairs 100j to 100k).
[0264] FIG. 26 is a timing diagram according to an example. The timing diagram of FIG. 26 is used as a reference as part of the process for explaining the asynchronous two-way ranging method. The symbols, acronyms, and their meanings used in this discussion are as follows:
[0265]
Table 1
Table 2
[0266] Furthermore, the acronym "DW" indicates a delayed waveform. The hat above the symbol indicates an estimated value. The tilde above the symbol indicates a measured value. The "clock epoch" of an audio device is the time when the audio device control system transmits the playback signal to the loudspeaker(s) or loudspeaker(s). The "playback epoch" of an audio device is the time when the loudspeaker(s) or loudspeaker(s) actually plays the sound corresponding to the playback signal. The terms "latency" and "delay" are used synonymously. For example, "playback latency" is the delay between the time when the audio device control system transmits the playback signal to the loudspeaker(s) or loudspeaker(s) and the time when the loudspeaker(s) or loudspeaker(s) actually plays the sound corresponding to the playback signal. Similarly, "recording latency" is the delay between the time when the microphone receives the signal and the time when that signal is received by the control system.
[0267] Figure 26 shows the timing for estimating the playback and recording latency of audio device i. Assuming that the playback and recording input / output (I / O) streams are synchronized, the full-duplex audio thread is used as the audio device's clock. When outputting a signal in sync with JPEG0007860985000010.jpg73, playback latency JPEG0007860985000011.jpg74 JPEG0007860985000012.jpg74 i Until then, the signal will not be played from the speaker. In other words,
[0268]
number
[0269] Subsequently, there is an acoustic delay caused by the distance between the speaker and microphone on the audio device. The signal reaches the microphone of the same audio device via JPEG0007860985000014.jpg74. The received signal has recording latency before it enters the audio thread of the audio device. Only JPEG0007860985000015.jpg74 will be further delayed.
[0270]
number
[0271] The DW generated by the audio device has a delay. JPEG0007860985000017.jpg72 ii It has a peak located at , where, ~ This indicates the measured value. JPEG0007860985000018.jpg72 ii This represents the measured pseudo-range between audio device i and itself. The difference in code phase between the local replica generated by the audio thread and the signal in the microphone feed is the code delay of the peak DW.
number
[0272] Figure 27 is a timing diagram showing the relevant clock terms when estimating the time of flight between two asynchronous audio devices in one example. Here, we consider the case where both audio devices are reproducing an acoustic DSSS signal, and the DW is generated by processing the acoustic DSSS signal of the other audio device. As a result, the delay measurement... JPEG0007860985000020.jpg74 and The resulting image, JPEG0007860985000021.jpg74, corresponds to ToF between audio devices. Figure 27 shows the transmission from device i and reception at device j, and vice versa.
[0273] In this example, the symbols and acronyms in Figure 27 have the following meanings and contexts:
[0274] JPEG0007860985000022.jpg720 is synchronized with the audio threads on devices i and j, respectively. The actual acoustic delay between the two devices is the same. The image is JPEG0007860985000023.jpg727, and each acoustic path is indicated by green and blue arrows. • The code phase of the transmitted signal is such that at transmission time (ToT), the speaker of device i is ( This is JPEG0007860985000024.jpg715. • After ToF, this signal arrives at the receiving end (device j), but is delayed by the recording latency on device j. Therefore, the phase of the transmitted signal in the microphone buffer of the audio thread of device j is at the time of reception (ToR). This will result in JPEG0007860985000025.jpg724. • The code phase of the local replica generated by the audio thread running on device j is in ToR. It has the phase of JPEG0007860985000026.jpg722.
[0275] Since the difference in code phase between the local replica and the received signal determines where the DW peak occurs, the measured delay can be expressed as follows:
number
[0276] Performing a similar analysis, and determining the measured delay when device j transmits and i receives, we obtain the following equation:
number
[0277] See (5) and (6), relative clock bias term
number
number
[0278] Substituting (4) into (8) and rearranging, we obtain the following equation:
number
number
[0279] Therefore, using (9), an unbiased set of pseudorange estimates can be obtained if the following is obtained: • Mutual delay measurement: JPEG0007860985000033.jpg718 • Playback and recording latency measurements: JPEG0007860985000034.jpg718 • Estimated acoustic delay that constitutes the playback / recording latency of a specific device: JPEG0007860985000035.jpg718
[0280] In some examples, δ aThere are cases where it is neither possible to estimate nor eliminate δ in (9). a You can also omit this and retain the bias in the estimated pseudo-range:
number
[0281] Alternatively, δ based on the type of audio device a Use an approximate value of δ a You can also rely on measuring it in advance.
[0282] Clock bias estimate Instead of summing any two mutual pseudorange estimates, taking the difference yields the following:
number
number
number
[0283] Equation (14) allows us to solve for the relative clock bias (e.g., in a control system) Δ tij It is possible to solve this: 1. The difference between playback latency and recording latency is known (i.e., measured in advance and substituted into (14)), or 2. The difference between playback latency and recording latency is equal on both devices (the terms in (14) cancel each other out), or 3. The difference between playback latency and recording latency is zero (the terms cancel each other out in (13)).
[0284] Clock skew estimation Depending on the signal used to generate the DW, it may also be possible to process it in a way that allows us to obtain an estimate of the frequency difference (skew) between the clocks of two audio devices. The DSSS signal used in this experiment is a carrier signal located at f0Hz spread with a pseudo-random sequence (sometimes referred to herein as a PRN sequence, PRN code, spreading code, or simply a code). Receiving this signal involves both "de-spreading" and shifting back to the baseband. However, if the frequencies of the two clocks are different, after coherent integration (matched filtering using local replicas), there will be a residual frequency equal to the difference between the two clock frequencies. Therefore, instead of generating the DW by averaging the squares of the results of the coherent integration, in some forms it involves performing spectral analysis to determine the frequency of the residual carrier and inferring the difference in clock frequencies from the residual carrier frequency. In this way, the control system can obtain an estimate after one coherent integration period. However, unless the DSSS parameter set is modified to optimize for such measurements, the estimate is likely to be quite noisy even after only one coherent integration period. Such DSSS parameter changes may involve making the spreading code (and coherent integration period) very long in time (e.g., ranging from hundreds of milliseconds to several seconds), which can be achieved by using a longer code (more chips) and / or by reducing the chipping rate (bandwidth).
[0285] Another approach utilizes the fact that the relative code phase (and clock bias) walks (in other words, changes over time) due to the difference in clock frequencies. In some such modes, a control system can track how ~τij changes over time (this is the rate at which the code phase walks).
[0286] There is a trade-off between these two methods, which can be summarized as follows: • Carrier-based methods require spectral analysis for each coherent integral result, which introduces considerable complexity. Code walk-based methods require the control system to only process this amount of data by maintaining a history of the measured pseudorange sets, which is significantly smaller. If the clock frequency difference is large enough to be detected on the scale of the coherent integral period, there is a high probability of SNR loss in the DW, requiring a shorter period, which in turn makes it impossible to resolve the clock rate difference. Carrier-based methods generate estimates after only one coherent integration period, while code-walk-based methods require a sufficient number of DWs and pseudo-range estimates to reliably estimate code walks amid the phase noise of the DWs. Therefore, code-walk-based methods are significantly slower. However, coherent carrier-based methods, which are inherently noisy, may require temporal smoothing, potentially resulting in similar observation times.
[0287] In some cases, a delay rate estimator (for example, as described above, see Figure 24) can be used to estimate the clock skew. The delay rate is proportional to the clock skew.
[0288] Figure 28 is a graph illustrating an example of detecting relative clock skew between two audio devices using a single acoustic DSSS signal. In this example, the horizontal axis represents frequency and the vertical axis represents power. Figure 28 shows the spectrum of the main lobe of the received modulated acoustic DSSS signal 2807 and the frequency of the demodulated acoustic DSSS signal 2808. Note that the demodulated acoustic DSSS signal 2808 is not zero Hz. This represents the relative clock skew between the devices.
[0289] Figure 29 is a graph illustrating an example of detecting relative clock skew between two audio devices by measuring a single acoustic DSSS signal multiple times. In this example, the horizontal axis represents delay time, and the vertical axis represents power. Figure 29 shows an example of a set of delay waveforms generated from a set of acoustic DSSS signals for blocks of received audio (t=1 and t=2). The shift in the position of the delay waveform peaks (which itself indicates bulk delay) indicates clock skew between devices. In some examples, time 2 may be several hours or even days after time 1. Using such relatively large time intervals is advantageous when the clock skew is relatively small.
[0290] Clock discipline In some forms, control systems are configured to actually drive (discipline) a local clock using closed-loop methods, leveraging clock bias and delay estimation. A signal processing chain can be implemented to achieve clock discipline using frequency-locked loops, delay-locked loops, phase-locked loops, or a combination thereof.
[0291] In an alternative example, instead of actually adjusting the local clock, the DSSS signal parameters may be adjusted to compensate for the clock bias.
[0292] Since the accuracy of clock bias and delay estimation techniques is highly dependent on SNR, (see Figure 7) the optimization module 712 would be best suited for determining the DSSS parameter set 705 by giving relatively higher weight to acoustic DSSS signal performance estimates (single or multiple) 703 than to perceptual effect estimates (single or multiple) 702. For example, the optimization module 712 may be configured to determine the DSSS parameter set 705 by emphasizing the system's ability to generate a high SNR observation set of acoustic DSSS signals, and not emphasizing the impact / perceptibility of the acoustic DSSS signals to the user. In some such examples, the DSSS parameter set 705 can correspond to an audible acoustic DSSS signal set.
[0293] However, in some alternatives, coarse techniques (such as DW delayed tracking) may be implemented in a way that is sub-audible and has a low signal-to-noise ratio.
[0294] Device discoverability Figure 30 is a graph illustrating an example of an acoustic DSSS spreading code reserved for device discovery. In this example, the reserved spreading code is used, for example, when a new audio device is powered on and in the process of being configured for use in an audio environment. During runtime operation, a different ("normal") acoustic DSSS spreading code is used. The reserved spreading code may or may not use the same frequency band as the normal acoustic DSSS spreading code.
[0295] The elements of Figure 30 are as follows: 3001: A set of reserved acoustic DSSS spreading codes, also known as pseudorandom number sequences. 3002: Multiple pseudo-random number sequences assigned (by the orchestration device). 3003: Device 1 already has a code assigned to it. 3006: Device 2 is sending reservation code (3001). 3004: Device 2 is detected, and the orchestration device assigns a code to device 2. 3007: Device 2 is sending the code assigned to it. 3008: Device 3 begins transmitting the reservation code after it is powered on for the first time. 3005: Device 3 is detected, and the orchestration device assigns a code to device 3. 3009: Device 3 is sending the code assigned to it.
[0296] In this example, when a new audio device is introduced into the audio environment system, the new audio device begins to reproduce the generated acoustic DSSS signal using a reserved spread code sequence. This allows other devices in the room to recognize that the new audio device has been introduced into the acoustic space, and the integration sequence begins. After the new audio device is discovered and integrated into the system of the orchestrated audio devices, the new audio device begins to reproduce the acoustic DSSS signal sequence using a spread code (assigned by the orchestration device in this example).
[0297] In this example, devices 2 and 3 move from the discovery code channel (frequency band) to the frequency band assigned by the orchestration system. During integration, the amplitude, bandwidth, and center frequency of all devices reproducing the acoustic DSSS signal set may be modified to ensure optimal observation for the new system configuration. In some examples, the orchestration device may recalculate the acoustic DSSS parameter set for all devices in the acoustic space, so a newly discovered audio device may change the DSSS parameter set for all audio devices.
[0298] Noise estimation In this example, a set of acoustic DSSS-based observations generated by multiple audio devices is used to estimate the noise in the acoustic space.
[0299] Figure 31 shows another example of an audio environment. In Figure 31, an acoustic space 130 is shown having multiple distributed orchestrated audio devices 100A, 100B, and 100C that participate in DSSS operation. In this example, there is also a noise source 8500 that generates noise 8501. The elements of Figure 31 are as follows:
[0300] 130: Acoustic space; 100(A,B,C): Multiple distributed audio devices undergoing orchestration; 110: Multiple loudspeakers; 111: Multiple microphones; 8010: Distance between 100A and 100B; The distance between 8011:100A and 100C; The distance between 8012:100B and 100C; 8500: Noise source; 8501: Noise; The distance between 8510, 8500, and 100A; The distance between 8511:8500 and 100B; and The distance between 8512, 8500, and 100C.
[0301] Figure 32A shows an example of a set of delay waveforms generated by audio device 100C in Figure 31, based on the acoustic DSSS signals received from audio devices 100A and 100B. The delay waveform corresponding to the acoustic DSSS signals received from audio device 100A is shown as 400Ca, and the delay waveform corresponding to the acoustic DSSS signals received from audio device 100B is shown as 400Cb.
[0302] Figure 32B shows an example of a group of delay waveforms generated by audio device 100B in Figure 31 based on the acoustic DSSS signal groups received from audio devices 100A and 100C. The delay waveform corresponding to the acoustic DSSS signal group received from audio device 100A is shown as 400Ba, and the delay waveform corresponding to the acoustic DSSS signal group received from audio device 100C is shown as 400Bc.
[0303] The elements of Figures 32A and 32B are as follows: 400Ca: Delay waveform generated by device 100C in response to the acoustic DSSS signal group received from 100A; Delay waveform generated by device 100C in response to the acoustic DSSS signal group received from 400Cb:100B; 400Ba: Delay waveform generated by device 100B in response to the acoustic DSSS signal group received from 100A; 400Bc: Delay waveform generated by device 100B in response to the acoustic DSSS signal group received from 100C; 401C, 401B: Noise floor region of the delayed waveform group; 8552Ca: The signal power of the delayed waveform generated by 100C in response to the acoustic DSSS signal group received from 100A; 8552Cb: The signal power of the delayed waveform generated by 100C in response to the acoustic DSSS signal group received from 100B; 8552Ba: Signal power of the delayed waveform generated by 100B in response to the acoustic DSSS signal group received from 100A; 8552Bc: The signal power of the delayed waveform generated by 100B in response to the acoustic DSSS signal group received from 100C; 8551Ca: Noise power of the delayed waveform generated by 100C in response to the acoustic DSSS signal group received from 100A; 8551Cb: The noise power of the delayed waveform generated by 100C in response to the acoustic DSSS signal group received from 100B; 8551Ba: Noise power of the delayed waveform generated by 100B in response to the acoustic DSSS signal group received from 100A; and 8551Bc: Noise power of the delayed waveform generated by 100B in response to the acoustic DSSS signal group received from 100C.
[0304] Referring again to Figure 31, in this example, the distance 8511 between audio device 100B and noise source 8500 is shorter than the distance 8512 between audio device 100C and noise source 8500, and also shorter than the distance 8510 between audio device 100A and noise source 8500. In this particular scenario, the relative proximity of audio device 100B and noise source 8500 results in noise powers 8551Ba and 8551Bc in signals 400Ba and 400Bc being greater than noise powers 8551Ca and 8551Cb in signals 400Ca and 400Cb. Furthermore, signal 400Bc has relatively more noise than signal 400Ba. This suggests that noise source 8500 is located closer to the path between audio devices 100B and 100C than to the path between audio devices 100B and 100A. In some configurations, one or more audio devices may include a directional microphone or be configured for receiver beamforming. Such capability can provide further information about the DoA of sound from a noise source, and therefore further information about the location of the noise source.
[0305] Therefore, using the known or calculated locations of audio devices, the known or calculated distances between audio devices, the measured locations of noise sources, and the relative noise levels of the delayed waveform sets generated by each audio device, in some examples, the control system may be configured to generate distributed noise estimates for the audio environment 130. Such distributed noise estimates may be, or be based on, a set of noise estimates measured by microphones on audio devices located at different locations in the acoustic space. For example, one audio device might be placed near a kitchen bench, another near a lounge chair, and yet another near a door. Each of these devices is more sensitive to noise in its immediate vicinity and to various locations in the acoustic space, and together they can generate a set of estimates for the noise distribution of the entire room. Some such embodiments may involve the control system applying an assumed attenuation function based on the distance between the group of audio devices and the noise source. Some such examples may involve the control system comparing the calculated noise levels at each audio device to the measured noise floor of the delayed waveform set and / or the difference between the measured noise floors of the delayed waveform set (e.g., the difference in level or power between 8551Ca and 8551Cb).
[0306] Figure 33 is a flowchart illustrating another example of the disclosed method. The blocks of Method 3300, as with other methods described herein, are not necessarily performed in the order shown. Also, such a method may include more or fewer blocks than those illustrated and / or described. Method 3300 may be performed by an apparatus or system such as the apparatus 150 shown in Figure 1B and described above.
[0307] In this example, block 3305 includes receiving a first content stream containing a first set of audio signals by a control system. The content stream and the first set of audio signals may vary depending on a particular mode. In some examples, the content stream may correspond to television programs, movies, music, podcasts, etc.
[0308] In this example, block 3310 includes the control system rendering the first audio signal set to generate a first audio playback signal set. The first audio playback signal set may be, or include, a loudspeaker feed signal set for the loudspeaker system of an audio device.
[0309] In this example, block 3315 includes generating a first direct sequence spread spectrum (DSSS) signal set by the control system. In this example, the first DSSS signal set corresponds to a signal referred to herein as an acoustic DSSS signal. In some examples, the first DSSS signal set may be generated by one or more DSSS signal generator modules, such as the DSSS signal generator 212A and DSSS signal modulator 220A described above with reference to Figure 2.
[0310] In this example, block 3320 includes generating a first modified audio playback signal group by inserting the first DSSS signal group into the first audio playback signal group by the control system. In some examples, block 3320 may be performed by the DSSS signal injector 211A described above with reference to Figure 2.
[0311] In this example, block 3325 includes generating a first audio device playback sound by causing the loudspeaker system to reproduce the first modified audio playback signal group by the control system. In some examples, block 3320 may include controlling the loudspeaker system 110A so that the control system 160 in Figure 2 reproduces the first modified audio playback signal group and generates a first audio device playback sound.
[0312] In some embodiments, method 3300 may include the control system receiving a group of microphone signals from the microphone system corresponding to at least the first audio device playback sound and the second audio device playback sound. The second audio device playback sound may correspond to a second group of modified audio playback signals reproduced by the second audio device. In some examples, the second group of modified audio playback signals may include a second group of DSSS signals generated by the second audio device. In some such examples, method 3300 may include the control system extracting at least the second group of DSSS signals from the group of microphone signals.
[0313] In some embodiments, method 3300 may include the control system receiving a group of microphone signals from the microphone system corresponding to at least the first audio device playback sound and the second to Nth audio device playback sounds. The second to Nth audio device playback sounds may correspond to the second to Nth modified audio playback signals reproduced by the second to Nth audio devices. In some examples, the second to Nth modified audio playback signals may include the second to Nth DSSS signals. In some such examples, method 3300 may include the control system extracting at least the second to Nth DSSS signals from the group of microphone signals.
[0314] In some embodiments, Method 3300 may include the control system estimating at least one acoustic scene metric based at least partially on the second to nth DSSS signals. In some examples, the acoustic scene metric(s) may be or include time of flight, time of arrival, range, audibility of the audio device, impulse response of the audio device, angle between audio devices, position of the audio device, noise of the audio environment, and / or signal-to-noise ratio. According to some examples, Method 3300 may include the control system controlling one or more modes of audio device playback based at least partially on at least one acoustic scene metric and / or at least one audio device characteristic.
[0315] In some examples, the first content stream component of the sound played by the first audio device may cause perceptual masking of the first DSSS signal component of the sound played by the first audio device. In some such examples, the first DSSS signal component may be inaudible to humans.
[0316] In some examples, method 3300 may include the control system determining one or more DSSS parameters for each of the multiple audio devices in the audio environment. The one or more DSSS parameters may be available for generating a set of DSSS signals. Some such examples may include the control system providing the one or more DSSS parameters to each of the multiple audio devices.
[0317] In some embodiments, determining the one or more DSSS parameters may include scheduling a time slot for each of the plurality of audio devices to reproduce the modified audio playback signal set. In some such examples, the first time slot for the first audio device may be different from the second time slot for the second audio device.
[0318] In some examples, determining the one or more DSSS parameters may include determining the frequency band for reproducing the modified audio playback signal set for each of the plurality of audio devices. In some such examples, the first frequency band for the first audio device may be different from the second frequency band for the second audio device.
[0319] In some examples, determining the one or more DSSS parameters may include determining a DSSS spreading code for each of the multiple audio devices. In some such examples, a first spreading code for a first audio device may be different from a second spreading code for a second audio device. In some examples, determining the one or more DSSS parameters may include determining at least one spreading code length, which is at least partially based on the corresponding audio device audibility. According to some examples, determining the one or more DSSS parameters may include applying an acoustic model which is at least partially based on the interaudibility of each of the multiple audio devices in the audio environment.
[0320] In some examples, determining the one or more DSSS parameters may include determining the current playback objective. Some such examples may include determining the estimated performance of the DSSS signal set in the audio environment by applying an acoustic model that is at least partially based on the interaudibility of each of the multiple audio devices in the audio environment. Some such examples may include determining the perceptual effect of the DSSS signal set in the audio environment by applying a perceptual model that is based on human auditory perception. Some such examples may include determining the one or more DSSS parameters at least partially based on the current playback objective, the estimated performance, and / or the perceptual effect.
[0321] According to some examples, determining the one or more DSSS parameters may include detecting a DSSS parameter change trigger and determining one or more new DSSS parameters corresponding to the DSSS parameter change trigger. Some such examples may include providing the one or more new DSSS parameters to one or more audio devices in the audio environment.
[0322] In some examples, detecting the DSSS parameter change trigger may include detecting one or more of the following: a new audio device in the audio environment, a change in the position of an audio device, a change in the orientation of an audio device, a change in the configuration of an audio device, a change in the position of a person in the audio environment, a change in the type of audio content being played in the audio environment, a change in background noise in the audio environment, a change in the configuration of the audio environment including but not limited to a change in the configuration of doors or windows in the audio environment, a clock skew between two or more audio devices in the audio environment, a clock bias between two or more audio devices in the audio environment, a change in the interacousability between two or more audio devices in the audio environment, and / or a change in the purpose of playback.
[0323] In some examples, method 3300 may include processing the received microphone signal set to generate a preprocessed microphone signal set. Some such examples may include extracting a DSSS signal set from the preprocessed microphone signal set. Processing the received microphone signal set may include, for example, applying beamforming, bandpass filtering, and / or echo cancellation.
[0324] In some embodiments, extracting at least the second to nth DSSS signals from the microphone signal group may include applying a matched filter to the microphone signal group or a pre-processed version of the microphone signal group to generate the second to nth delay waveforms. The second to nth delay waveforms may correspond, for example, to each of the second to nth DSSS signals. Some such examples may include applying a low-pass filter to each of the second to nth delay waveforms.
[0325] In some examples, method 3300 may include implementing a demodulator via the control system. Some such examples may include applying the matching filter as part of the demodulation process performed by the demodulator. In some such examples, the output of the demodulation process may be a demodulated coherent baseband signal. Some examples may include estimating a bulk delay via the control system and providing the bulk delay estimate to the demodulator.
[0326] In some examples, Method 3300 may include implementing a baseband processor configured for baseband processing of the demodulated coherent baseband signals via the control system. In some such examples, the baseband processor may be configured to output at least one estimated acoustic scene metric. In some examples, the baseband processing may include generating a non-coherently integrated delayed waveform based on a group of demodulated coherent baseband signals received during a non-coherent integration period. In some such examples, generating the non-coherently integrated delayed waveform may include squaring the group of demodulated coherent baseband signals received during the non-coherent integration period to generate a group of squared demodulated baseband signals, and integrating the group of squared demodulated baseband signals. In some examples, the baseband processing may include applying one or more of the leading-edge estimation process, staired response power estimation process, or signal-to-noise estimation process to the non-coherently integrated delayed waveform. Some examples may include estimating the bulk delay via the control system and providing the bulk delay estimate to the baseband processor.
[0327] According to some examples, method 3300 may include the control system estimating second to nth noise power levels at the locations of second to nth audio devices based on the second to nth delayed waveforms. Some such examples may include generating a distributed noise estimate for the audio environment based at least in part on the second to nth noise power levels.
[0328] In some examples, method 3300 may include performing an asynchronous bidirectional ranging process to offset an unknown clock bias between two asynchronous audio devices. The asynchronous bidirectional ranging process may be based, for example, on a set of DSSS signals transmitted by each of the two asynchronous audio devices. Some such examples may include performing the asynchronous bidirectional ranging process between each of a plurality of audio device pairs in the audio environment.
[0329] According to some examples, method 3300 may include performing a clock bias estimation process to determine an estimated clock bias between two asynchronous audio devices. The clock bias estimation process may be based, for example, on a set of DSSS signals transmitted by each of the two asynchronous audio devices. Some such examples may include compensating for the estimated clock bias. Some embodiments may include performing the clock bias estimation process between each of a plurality of audio devices in the audio environment to generate a plurality of estimated clock biases. Some such embodiments may include compensating for each estimated clock bias.
[0330] In some examples, method 3300 may include performing a clock skew estimation process to determine estimated clock skew between two asynchronous audio devices. The clock skew estimation process may be based, for example, on a set of DSSS signals transmitted by each of the two asynchronous audio devices. Some such examples may include compensating for the estimated clock skew. Some such examples may include performing the clock skew estimation process between each of a plurality of audio device pairs in the audio environment to generate a plurality of estimated clock skews. Some such examples may include compensating for each estimated clock skew.
[0331] According to some examples, method 3300 may include detecting a DSSS signal transmitted by an audio device. In some examples, the DSSS signal may correspond to a first spreading code. Some such examples may include providing a second spreading code for later transmission to the audio device. In some such examples, the first spreading code may include a first pseudorandom number sequence reserved for the newly activated audio device.
[0332] In some examples, method 3300 may include simultaneously playing a modified audio playback signal set on each of the multiple audio devices in the audio environment.
[0333] In some examples, the acoustic DSSS signal set may be reproduced over one or more time intervals in which the audio playback signal set is inaudible, referred herein as “silent intervals” or “silence.” In some such examples, at least a portion of the first audio signal set may correspond to silence.
[0334] Figure 34 is a flowchart illustrating another example of the disclosed method. The blocks of Method 3400, as with other methods described herein, are not necessarily performed in the order shown. Also, such a method may include more or fewer blocks than those illustrated and / or described. Method 3400 may be performed by an apparatus or system such as Apparatus 150 shown in Figure 1B and described above.
[0335] In some examples, the blocks of Method 3400 may be implemented by one or more devices in an audio environment, such as an orchestration device, such as an audio system controller (referred to herein as a smart home hub), or by another component of the audio system, such as a smart speaker, television, television control module, laptop computer, or mobile device (such as a mobile phone). In some embodiments, the audio environment may include one or more rooms in a home environment. In other embodiments, the audio environment may be another type of environment, such as an office environment, a car environment, a train environment, a road or sidewalk environment, or a park environment. However, in other embodiments, at least some blocks of Method 3400 may be implemented by a device that implements cloud-based services, such as a server.
[0336] In this example, block 3405 includes a control system causing a first audio device in an audio environment to generate a first direct sequence spread spectrum (DSSS) signal set. In this example, the first DSSS signal set corresponds to a signal referred to herein as an acoustic DSSS signal. In some examples, the first DSSS signal set may be generated by one or more DSSS signal generator modules, such as the DSSS signal generator 212A and DSSS signal modulator 220A described above with reference to Figure 2, in accordance with instructions received from an orchestration device. Thus, the control system may be an orchestration device control system. In some examples, the instructions may be received from an orchestration module of an audio device, for example, from the orchestration module of the first audio device.
[0337] In this example, block 3410 includes causing the control system to insert the first DSSS signal set into a first audio playback signal set corresponding to a first content stream, thereby generating a first modified audio playback signal set for the first audio device. In some examples, block 3410 may be executed by the DSSS signal injector 211A described above with reference to Figure 2, in accordance with instructions received from an orchestration device or orchestration module.
[0338] In this example, block 3415 includes the control system causing the first audio device to reproduce the first modified audio playback signal set and generate the first audio device playback sound. In some examples, block 3415 may include the control system 160 in Figure 2 controlling the loudspeaker system 110A (according to commands received from an orchestration device or orchestration module) to reproduce the first modified audio playback signal set and generate the first audio device playback sound.
[0339] In some embodiments, blocks 3405, 3410, and 3415 may include providing DSSS information (such as DSSS information 205A as described above with reference to Figure 2) to the first audio device in the audio environment via an orchestration device or orchestration module. As described above, the DSSS information may include a set of parameters used by the control system of the first audio device for generating a set of DSSS signals, modulating a set of DSSS signals, demodulating a set of DSSS signals, and so on. The DSSS information may include one or more DSSS spreading code parameters and one or more DSSS carrier parameters, for example, as described elsewhere in this specification.
[0340] In this example, block 3420 includes causing the control system to cause a second audio device in the audio environment to generate a second DSSS signal set. In this embodiment, block 3425 includes causing the control system to insert the second DSSS signal set into a second content stream to generate a second modified audio playback signal set for the second audio device. In this embodiment, block 3430 includes causing the control system to cause the second audio device to play the second modified audio playback signal set to generate a second audio device playback sound. Blocks 3420-3430 may be executed, for example, in accordance with blocks 3405-3415. In some embodiments, 3420-3430 are executed in parallel with blocks 3405-3415.
[0341] In this example, block 3435 includes causing the control system to cause at least one microphone in the audio environment to detect at least the first audio device playback sound and the second audio device playback sound, and to generate a group of microphone signals corresponding to at least the first audio device playback sound and the second audio device playback sound. The at least one microphone may be a component of one or more audio devices in the audio environment, such as the first audio device, the second audio device, another audio device (such as the orchestration device), or other audio devices.
[0342] In this example, block 3440 includes causing the control system to extract the first DSSS signal group and the second DSSS signal group from the microphone signal group. Block 3440 may be performed by one or more audio devices in the audio environment, for example, having at least one microphone as mentioned in block 3435.
[0343] In this example, block 3445 includes causing the control system to estimate at least one acoustic scene metric based at least partially on the first DSSS signal set and the second DSSS signal set. The at least one acoustic scene metric may include, for example, one or more of the following: time of flight, time of arrival, range, audibility of an audio device, impulse response of an audio device, angle between audio devices, position of an audio device, noise of the audio environment, or signal-to-noise ratio.
[0344] In some examples, estimating the at least one acoustic scene metric may include estimating the at least one acoustic scene metric or having another device estimate the at least one acoustic scene metric. That is, the acoustic scene metric may be estimated by an orchestration device or another device in the audio environment.
[0345] In some embodiments, method 3400 may include controlling one or more aspects of audio device playback based at least partially on the at least one acoustic scene metric. For example, some embodiments may include controlling a noise compensation process based at least partially on one or more acoustic scene metrics. Some examples may include controlling a rendering process and / or one or more audio device playback levels based at least partially on one or more acoustic scene metrics.
[0346] In some manner, the DSSS signal component of the audio device playback sound may not be audible to humans. In some examples, the first content stream component of the first audio device playback sound may cause perceptual masking of the first DSSS signal component of the first audio device playback sound. In some examples, the second content stream component of the second audio device playback sound may cause perceptual masking of the second DSSS signal component of the second audio device playback sound.
[0347] In some examples, Method 3400 may include a control system causing a third to nth audio device in the audio environment to generate a third to nth direct sequence spread spectrum (DSSS) signal. Some such examples may include the control system generating a third to nth modified audio playback signal for the third to nth audio device by inserting the third to nth DSSS signals into the third to nth content streams. Some such examples may include the control system generating a third to nth instance of the audio device playback sound by causing the third to nth audio device to play a corresponding instance of the third to nth modified audio playback signal.
[0348] In some examples, method 3400 may include simultaneously playing a modified audio playback signal set on each of the multiple audio devices in the audio environment.
[0349] Some such examples may include the control system causing at least one microphone of each of the first to nth audio devices to detect first to nth instances of the audio device playback sound, and generating a set of microphone signals corresponding to the first to nth instances of the audio device playback sound. In some such examples, the first to nth instances of the audio device playback sound may include the first audio device playback sound, the second audio device playback sound, and the third to nth instances of the audio device playback sound. Some such examples may include the control system causing the first to nth DSSS signals to be extracted from the set of microphone signals, wherein the at least one acoustic scene metric is estimated at least in part based on the first to nth DSSS signals.
[0350] In some examples, method 3400 may include determining one or more DSSS parameters for a plurality of audio devices in the audio environment. The one or more DSSS parameters may be available for generating a set of DSSS signals. Some such examples may include providing the one or more DSSS parameters to each audio device of the plurality of audio devices. In some examples, determining the one or more DSSS parameters may include scheduling a time slot for each audio device of the plurality of audio devices to reproduce a set of modified audio playback signals. In some examples, a first time slot for a first audio device may be different from a second time slot for a second audio device.
[0351] In some examples, determining the one or more DSSS parameters may include determining the frequency band for reproducing the modified audio playback signal set for each of the plurality of audio devices. In some examples, the first frequency band for the first audio device may be different from the second frequency band for the second audio device.
[0352] In some examples, determining the one or more DSSS parameters may include determining a spreading code for each of the multiple audio devices. In some examples, a first spreading code for a first audio device may be different from a second spreading code for a second audio device. In some examples, determining the one or more DSSS parameters may include determining at least one spreading code length, which is at least partially based on the audibility of the corresponding audio device.
[0353] According to some examples, determining the one or more DSSS parameters may involve applying an acoustic model based at least in part on the interaudibility of each of the multiple audio devices in the audio environment.
[0354] In some examples, determining the one or more DSSS parameters may include determining the current playback objective. Some such examples may include determining the estimated performance of the DSSS signal set in the audio environment by applying an acoustic model that is at least partially based on the interaudibility of each of the multiple audio devices in the audio environment. Some such examples may include determining the perceptual effect of the DSSS signal set in the audio environment by applying a perceptual model that is based on human auditory perception. Some such examples may include determining the one or more DSSS parameters at least partially based on the current playback objective, the estimated performance, and the perceptual effect.
[0355] According to some examples, determining the one or more DSSS parameters may include detecting a DSSS parameter change trigger. Some such examples may include determining one or more new DSSS parameters corresponding to a DSSS parameter change trigger. Some such examples may include providing the one or more new DSSS parameters to one or more audio devices in the audio environment.
[0356] In some examples, detecting the DSSS parameter change trigger may include detecting one or more of the following: a new audio device in the audio environment, a change in the position of an audio device, a change in the orientation of an audio device, a change in the configuration of an audio device, a change in the position of a person in the audio environment, a change in the type of audio content being played in the audio environment, a change in background noise in the audio environment, a change in the configuration of doors or windows in the audio environment, a clock skew between two or more audio devices in the audio environment, a clock bias between two or more audio devices in the audio environment, a change in the interacousability between two or more audio devices in the audio environment, and / or a change in the purpose of playback.
[0357] According to some examples, method 3400 may include processing the received microphone signal set to generate a preprocessed microphone signal set. In some such examples, the DSSS signal set may be extracted from the preprocessed microphone signal set. In some such examples, processing the received microphone signal set may include applying one or more of beamforming, bandpass filtering, or echo cancellation.
[0358] In some examples, extracting at least the first DSSS signal group and the second DSSS signal group from the microphone signal group may include applying a matched filter to the microphone signal group or a pre-processed version of the microphone signal group to generate a delay waveform group. In some such examples, the delay waveform group may include at least a first delay waveform based on the first DSSS signal group and a second delay waveform based on the second DSSS signal group. In some examples, this may include applying a low-pass filter to the delay waveform group.
[0359] In some examples, applying the matching filter is part of the demodulation process. In some such examples, the demodulation process may be performed by demodulator 214A described above with reference to Figure 2, demodulator 214 described above with reference to Figure 17, or demodulator 214 described above with reference to Figure 18. In some such examples, the output of the demodulation process may be a demodulated coherent baseband signal. In some examples, this may include estimating a bulk delay and providing the bulk delay estimate to the demodulation process.
[0360] Some examples may include performing baseband processing on the demodulated coherent baseband signal, for example, by an instance of the baseband processor 218 disclosed herein. In some examples, the baseband processing may output at least one estimated acoustic scene metric. In some examples, the baseband processing may include generating a non-coherently integrated delayed waveform based on a group of demodulated coherent baseband signals received during a non-coherent integration period. According to some such examples, generating the non-coherently integrated delayed waveform may include squaring the group of demodulated coherent baseband signals received during the non-coherent integration period to generate a group of squared demodulated baseband signals, and integrating the group of squared demodulated baseband signals. In some embodiments, the baseband processing may include applying a leading-edge estimation process, a staired response power estimation process, and / or a signal-to-noise estimation process to the non-coherently integrated delayed waveform. Some examples may include estimating the bulk delay and providing the bulk delay estimate to the baseband processing.
[0361] Some examples may include estimating at least a first noise power level at the location of a first audio device and estimating a second noise power level at the location of a second audio device. In some such examples, the estimation of the first noise power level may be based on the first delay waveform, and the estimation of the second noise power level may be based on the second delay waveform. Some such examples may include generating a distributed noise estimate for the audio environment based at least in part on the estimated first noise power level and the estimated second noise power level.
[0362] In some examples, method 3400 may include performing an asynchronous bidirectional ranging process to offset an unknown clock bias between two asynchronous audio devices. In some examples, the asynchronous bidirectional ranging process may be based on a set of DSSS signals transmitted by each of the two asynchronous audio devices. In some examples, the asynchronous bidirectional ranging process may include performing the process between each of a plurality of pairs of audio devices in the audio environment.
[0363] According to some examples, method 3400 may include performing a clock bias estimation process to determine an estimated clock bias between two asynchronous audio devices. In some examples, the clock bias estimation process may be based on a set of DSSS signals transmitted by each of the two asynchronous audio devices. Some such examples may include compensating for the estimated clock bias. Some embodiments may include performing the clock bias estimation process between each of a plurality of audio devices in the audio environment to generate a plurality of estimated clock biases. Some such examples may include compensating for each of the plurality of estimated clock biases.
[0364] In some examples, method 3400 may include performing a clock skew estimation process to determine estimated clock skew between two asynchronous audio devices. The clock skew estimation process may be based on a set of DSSS signals transmitted by each of the two asynchronous audio devices. Some such examples may include compensating for the estimated clock skew. Some examples may include performing the clock skew estimation process between each of a plurality of audio devices in the audio environment to generate a plurality of estimated clock skews. Some such examples may include compensating for each of the plurality of estimated clock skews.
[0365] In some examples, method 3400 may include detecting a DSSS signal transmitted by an audio device. In some examples, the DSSS signal may correspond to a first spreading code. In some such examples, the method may include providing a second spreading code to the audio device. In some examples, the first spreading code may be, or include, a first pseudorandom number sequence reserved for a newly activated audio device.
[0366] In some examples, the acoustic DSSS signal group may be reproduced during one or more time intervals in which the audio playback signal group is inaudible. In some such examples, at least a portion of the first audio playback signal group, at least a portion of the second audio playback signal group, or at least a portion of each of the first and second audio playback signal groups corresponds to silence.
[0367] Figures 35, 36A, and 36B are flowcharts illustrating examples of how multiple audio devices can coordinate a measurement session in several modes. The blocks shown in Figures 35–36B, as with other methods described herein, are not necessarily performed in the order shown. For example, in some modes, the operation of block 3501 in Figure 35 may be performed before the operation of block 3500. Furthermore, such methods may include more or fewer blocks than those illustrated and / or described.
[0368] In these examples, the smart audio device is the orchestration device (which may also be referred to herein as the “leader”), and only one device can be an orchestration device at a time. In other examples, the orchestration device may be what is referred to herein as the smart home hub. The orchestration device may be an instance of the apparatus 150 described above with reference to Figure 1B.
[0369] Figure 35 shows the blocks performed by all participating audio devices in this example. In this example, block 3500 includes obtaining a list of all other participating audio devices. The list in block 3500 can be created, for example, by aggregating information from other audio devices via network packets: other audio devices can, for example, broadcast their intention to participate in the measurement session. The list in block 3500 may be updated as audio devices are added to and / or removed from the audio environment. In some such examples, the list in block 3500 may be updated according to various heuristics to keep the list up-to-date with respect only to the most important devices (e.g., audio devices currently in the main living space 130 in Figure 1A).
[0370] In the example shown in Figure 35, link 3504 indicates that the list of block 3500 is passed to block 3501, the leadership negotiation process. This negotiation process in block 3501 can take different forms depending on the specific configuration. In the simplest embodiment, assuming all devices can implement the same scheme, a leader can be determined without multiple communication rounds between devices by alphanumeric sorting for the lowest or highest device ID code (or other unique device identifier). In more complex configurations, devices can negotiate with each other to determine which device is best suited to be the leader. For example, a device that aggregates organized information might also be a suitable leader for the purpose of facilitating a measurement session. The device with the longest uptime, the device with the highest computing power, and / or the device connected to the mains power might be a suitable candidate for leader. In general, coordinating such consensus among multiple devices is a difficult problem, but it is a problem for which there are many existing satisfactory protocols and solutions (e.g., the Paxos protocol). It will be understood that there are many such protocols, and they may be appropriate.
[0371] In this example, all participating audio devices then execute block 3503. This means that in this example, link 3506 is an unconditional link. Block 3503 will be discussed later with reference to Figure 36B. If a device is the leader, that device executes block 3502. In this example, link 3505 includes a leadership check. An example of the leadership process is described below with reference to Figure 36A. The output from this leadership process (including, but not limited to, messages to other audio devices) is shown by link 3507 in Figure 35.
[0372] Figure 36A shows an example of a process performed by an orchestration device or leader. Block 3601 includes determining a set of acoustic DSSS parameters for each participating audio device. In some examples, block 3601 may include determining one or more DSSS spreading code parameters and one or more DSSS carrier parameters. In some examples, block 3601 may include determining a spreading code for each participating audio device. According to some such examples, a first spreading code for a first audio device may be different from a second spreading code for a second audio device. In some examples, block 3601 may include determining a spreading code length based at least in part on the audibility of the corresponding audio device. According to some examples, block 3601 may be based at least in part on the current playback purpose. In some examples, block 3601 may be based at least in part on whether a DSSS parameter change trigger has been detected.
[0373] In this example, after the orchestration device determines the acoustic DSSS parameter set in block 3601, the process in Figure 36A continues to block 3602. In this example, block 3602 includes sending the acoustic DSSS parameter set determined in block 3601 to other participating audio devices. In some examples, block 3602 may include sending the acoustic DSSS parameter set to other participating audio devices via wireless communication, such as a local WiFi network or Bluetooth. In some examples, block 3602 may include sending a “start session” instruction, as described later with reference to Figure 36B. In some examples, participating audio devices update their own acoustic DSSS parameter set in block 502.
[0374] In this example, after block 3602, the process in Figure 36A continues to block 3603, where the orchestration device waits for the current measurement session to end. In this example, in block 3603, the orchestration device waits for confirmation that all other participating audio devices have ended their sessions. In other examples, block 503 may include waiting for a predetermined period of time. In some examples, block 503 may include waiting for a DSSS parameter change trigger to be detected.
[0375] In this example, after block 3603, the process in Figure 36A continues to block 3600, where the orchestration device is provided with information about the measurement session. Such information may influence the selection and timing of future measurement sessions. In some embodiments, block 3600 includes receiving a set of measurements obtained during the measurement session from all other participating audio devices. The type of received set of measurements may depend on the particular mode. According to some examples, the received set of measurements may be, or include, a set of microphone signals. Alternatively, or additionally, in some examples, the received set of measurements may be, or include, audio data extracted from the set of microphone signals. In some modes, the orchestration device may perform (or have performed) one or more operations on the received set of measurements. For example, the orchestration device may estimate (or have had estimated) the audibility of a target audio device or the location of a target audio device based at least in part on the extracted audio data. Some approaches may involve estimating the far-field audio environment impulse response and / or audio environment noise based at least partially on the extracted audio data.
[0376] In the example shown in Figure 36A, after block 3600 is executed, the process returns to block 3601. In some such examples, the process returns to block 3601 after a predetermined period of time following the execution of block 3600. In some examples, the process may return to block 3601 in response to user input. In some examples, the process may return to block 3601 after a DSSS parameter change trigger is detected.
[0377] Figure 36B shows an example of processing performed by participating audio devices other than the orchestration device. Here, block 3610 includes each of the other participating audio devices sending a transmission (e.g., a network packet) to the orchestration device to indicate that each device intends to participate in one or more measurement sessions. In some embodiments, block 3610 may also include sending the results of one or more previous measurement sessions to the leader.
[0378] In this example, block 3615 follows block 3610. According to this example, block 3615 includes waiting for notification that a new measurement session is being started, for example, via a “session start” packet.
[0379] In this example, block 3620 includes applying a set of DSSS parameters according to information provided by the orchestration device, along with, for example, a "session start" packet awaited in block 3615. In this example, block 3620 includes applying the DSSS parameters to generate a set of modified audio playback signals to be played back by the participating audio devices during the measurement session. In this example, block 3620 includes detecting the audio device playback sound via the audio device's microphone and generating a corresponding set of microphone signals during the measurement session. As suggested by link 3622, in some cases block 3620 may be repeated until all measurement sessions instructed by the orchestration device are completed (for example, according to a "stop" instruction (e.g., a stop packet) received from the orchestration device, or after a predetermined duration). In some examples, block 3620 may be repeated for each of a plurality of target audio devices.
[0380] Finally, block 3625 includes providing the orchestration device with the information obtained during the measurement session. In this example, after block 3625, the process in Figure 36B returns to block 3610. In some such examples, the process returns to block 3610 after a predetermined period of time following the execution of block 3625. In some examples, the process may return to block 3610 in response to user input.
[0381] Some aspects of this disclosure include a system or device (e.g., programmed) configured to perform one or more examples of the disclosed method, and a tangible computer-readable medium (e.g., a disk) for storing code for performing one or more examples of the disclosed method or its steps. For example, some disclosed systems are or include a programmable general-purpose processor, digital signal processor, or microprocessor, programmed and / or otherwise configured by software or firmware to perform any of a variety of operations on data, including embodiments of the disclosed method or its steps. Such a general-purpose processor is or includes a computer system comprising an input device, memory, and a processing subsystem programmed (and / or otherwise configured) to perform one or more examples of the disclosed method (or its steps) in response to data asserted therein.
[0382] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed or otherwise configured) to perform the necessary processing on one or more audio signals, including the execution of one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed system (or its elements) may be implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include input devices and memory) programmed and / or otherwise configured by software or firmware to perform any of the various operations including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the system of the present invention may be implemented as a general-purpose processor or DSP configured (e.g., programmed) to execute one or more examples of the disclosed methods, and the system may also include other elements (e.g., one or more loudspeakers and / or one or more microphones). The general-purpose processor configured to execute one or more examples of the disclosed methods may be coupled to input devices (e.g., a mouse and / or keyboard), memory, and a display device.
[0383] Another aspect of this disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) for storing code for performing the disclosed method or one or more examples of its steps (e.g., an executable coder).
[0384] While specific embodiments and applications of the Disclosure have been described herein, it will be apparent to those skilled in the art that many modifications are possible to the embodiments and applications described herein without departing from the scope of the disclosure described herein and claimed herein. Although specific forms of the Disclosure have been shown and described herein, it should be understood that the Disclosure is not limited to the specific embodiments or methods described herein.
Claims
1. The control system causes the first audio device in the audio environment to generate a first direct sequence spread spectrum (DSSS) signal group, The control system modulates the first DSSS signal group to generate a modulated first DSSS signal group. The control system causes the modulated first DSSS signal group to be inserted into the first audio playback signal group corresponding to the first content stream, thereby generating the first modified audio playback signal group for the first audio device. The control system causes the first audio device to reproduce the first modified audio playback signal group and generates the sound played back by the first audio device. The control system causes the second audio device of the audio environment to generate a second DSSS signal group, The control system modulates the second DSSS signal group to generate a modulated second DSSS signal group. The control system causes the modulated second DSSS signal group to be inserted into the second content stream, and generates a second modified audio playback signal group for the second audio device. The control system causes the second audio device to reproduce the second modified audio playback signal group, thereby generating the sound reproduced by the second audio device. The control system causes at least one microphone in the audio environment to detect at least the sound played by the first audio device and the sound played by the second audio device, and generates a group of microphone signals corresponding to at least the sound played by the first audio device and the sound played by the second audio device. The control system extracts the first DSSS signal group and the second DSSS signal group from the microphone signal group, The control system estimates at least one acoustic scene metric based at least partially on the first DSSS signal group and the second DSSS signal group, Controlling one or more modes of audio device playback based at least partially on the aforementioned at least one acoustic scene metric, Includes, An audio processing method wherein the at least one acoustic scene metric includes one or more of the following: time of flight, time of arrival, range, impulse response of an audio device, angle between audio devices, or position of an audio device.
2. The audio processing method according to claim 1, wherein the estimation of the at least one acoustic scene metric includes estimating the at least one acoustic scene metric or having another device estimate the at least one acoustic scene metric.
3. The audio processing method according to claim 1 or 2, wherein the first content stream component of the sound played back by the first audio device causes perceptual masking of the first DSSS signal component of the sound played back by the first audio device.
4. The audio processing method according to any one of claims 1 to 3, wherein the second content stream component of the sound played back by the second audio device causes perceptual masking of the second DSSS signal component of the sound played back by the second audio device.
5. The audio processing method according to any one of claims 1 to 4, wherein the control system is an orchestration device control system.
6. The control system causes the third to the Nth audio devices of the audio environment to generate the third to the Nth direct sequence spread spectrum (DSSS) signals, The control system generates modified audio playback signals for the third to nth audio devices by inserting the third to nth DSSS signals into the third to nth content streams, The control system causes the third to the nth audio devices to reproduce corresponding instances of the third to the nth modified audio playback signals, thereby generating third to nth instances of the audio device playback sound. The audio processing method according to any one of claims 1 to 5, further comprising:
7. The control system causes at least one microphone of each of the first to nth audio devices to detect first to nth instances of the audio device playback sound, and generates a group of microphone signals corresponding to the first to nth instances of the audio device playback sound, wherein the first to nth instances of the audio device playback sound include the first audio device playback sound, the second audio device playback sound, and the third to nth instances of the audio device playback sound. The control system causes the first to Nth DSSS signals to be extracted from the microphone signal group, and the at least one acoustic scene metric is estimated at least partially based on the first to Nth DSSS signals. The audio processing method according to claim 6, further comprising:
8. Determining one or more DSSS parameters for multiple audio devices in the aforementioned audio environment, wherein the one or more DSSS parameters can be used to generate a DSSS signal group including the first DSSS signal group and the second DSSS signal group. Providing one or more DSSS parameters to each of the multiple audio devices, The audio processing method according to any one of claims 1 to 7, further comprising:
9. The audio processing method according to claim 8, wherein determining one or more DSSS parameters includes scheduling a time slot for each of the plurality of audio devices to reproduce a group of modified audio playback signals, wherein the first time slot for the first audio device is different from the second time slot for the second audio device.
10. The audio processing method according to claim 8, wherein determining one or more DSSS parameters includes determining a frequency band for reproducing a modified audio playback signal group for each of the plurality of audio devices.
11. The audio processing method according to claim 10, wherein the first frequency band for the first audio device is different from the second frequency band for the second audio device.
12. The audio processing method according to any one of claims 8 to 10, wherein determining one or more DSSS parameters includes determining a spreading code for each of the plurality of audio devices.
13. The audio processing method according to claim 12, wherein the first spreading code for the first audio device is different from the second spreading code for the second audio device.
14. The audio processing method according to claim 12 or 13, further comprising determining at least one diffused code length based at least in part on the audibility of the corresponding audio device.
15. Determining one or more DSSS parameters means Detecting DSSS parameter change triggers, Determine one or more new DSSS parameters that correspond to the DSSS parameter change trigger, To provide the aforementioned one or more new DSSS parameters to one or more audio devices in the audio environment, An audio processing method according to any one of claims 8 to 14, including the following:
16. A control system causes a first audio device in an audio environment to generate a first direct sequence spread spectrum (DSSS) signal group, The control system modulates the first DSSS signal group to generate a modulated first DSSS signal group. The control system causes the modulated first DSSS signal group to be inserted into the first audio playback signal group corresponding to the first content stream, thereby generating the first modified audio playback signal group for the first audio device. The control system causes the first audio device to reproduce the first modified audio playback signal group and generates the sound played back by the first audio device. The control system causes the second audio device of the audio environment to generate a second DSSS signal group, The control system modulates the second DSSS signal group to generate a modulated second DSSS signal group. The control system causes the modulated second DSSS signal group to be inserted into the second content stream, and generates a second modified audio playback signal group for the second audio device. The control system causes the second audio device to reproduce the second modified audio playback signal group, thereby generating the sound reproduced by the second audio device. The control system causes at least one microphone in the audio environment to detect at least the sound played by the first audio device and the sound played by the second audio device, and generates a group of microphone signals corresponding to at least the sound played by the first audio device and the sound played by the second audio device. The control system extracts the first DSSS signal group and the second DSSS signal group from the microphone signal group, The control system estimates at least one acoustic scene metric based at least partially on the first DSSS signal group and the second DSSS signal group, Controlling one or more modes of audio device playback based at least partially on the aforementioned at least one acoustic scene metric, Determining one or more DSSS parameters for multiple audio devices in the aforementioned audio environment, wherein the one or more DSSS parameters can be used to generate a DSSS signal group including the first DSSS signal group and the second DSSS signal group. Providing one or more DSSS parameters to each of the multiple audio devices, Includes, An audio processing method comprising determining one or more DSSS parameters, which includes scheduling a time slot for each of the plurality of audio devices to reproduce a group of modified audio playback signals, wherein the first time slot for the first audio device is different from the second time slot for the second audio device.
17. A control system causes a first audio device in an audio environment to generate a first direct sequence spread spectrum (DSSS) signal group, The control system modulates the first DSSS signal group to generate a modulated first DSSS signal group. The control system causes the modulated first DSSS signal group to be inserted into the first audio playback signal group corresponding to the first content stream, thereby generating the first modified audio playback signal group for the first audio device. The control system causes the first audio device to reproduce the first modified audio playback signal group and generates the sound played back by the first audio device. The control system causes the second audio device of the audio environment to generate a second DSSS signal group, The control system modulates the second DSSS signal group to generate a modulated second DSSS signal group. The control system causes the modulated second DSSS signal group to be inserted into the second content stream, and generates a second modified audio playback signal group for the second audio device. The control system causes the second audio device to reproduce the second modified audio playback signal group, thereby generating the sound reproduced by the second audio device. The control system causes at least one microphone in the audio environment to detect at least the sound played by the first audio device and the sound played by the second audio device, and generates a group of microphone signals corresponding to at least the sound played by the first audio device and the sound played by the second audio device. The control system extracts the first DSSS signal group and the second DSSS signal group from the microphone signal group, The control system estimates at least one acoustic scene metric based at least partially on the first DSSS signal group and the second DSSS signal group, Controlling one or more modes of audio device playback based at least partially on the aforementioned at least one acoustic scene metric, Determining one or more DSSS parameters for multiple audio devices in the aforementioned audio environment, wherein the one or more DSSS parameters can be used to generate a DSSS signal group including the first DSSS signal group and the second DSSS signal group. Providing one or more DSSS parameters to each of the multiple audio devices, Includes, An audio processing method in which determining one or more DSSS parameters includes determining a frequency band for reproducing a modified audio playback signal group for each of the plurality of audio devices.
18. An apparatus configured to perform any one of the methods of claims 1 to 17.
19. A system configured to perform any one of the methods of claims 1 to 17.
20. A computer program comprising instructions for controlling one or more devices to perform the method according to any one of claims 1 to 17.